❌

Vue normale

Reçu avant avant-hiercloud

Celebrating our tech and startup customers

20 avril 2022 à 22:00

Our tech and startup customers are disrupting industries, driving innovation and changing how people do things. We’re proud of their success and want to showcase what they’re up to! You’ll hear about their new products, their businesses reaching new milestones and their ability to get things done faster and easier using Google Cloud’s app development, data analytics and AI/ML services.

Congrats to Impact Analytics for Closing PVH
With the COVID-19 pandemic, the rise of e-commerce, and supply chain crisis, Impact Analytics had to quickly offer enhancements on their platform that gave retailers access to intelligent, automated, and edge-aware solutions. Google Cloud's best in class AI and ML solutions and highly performant infrastructure gave Impact Analytics the scalability and building blocks to create Ada, a robust predictive algorithm to give PVH and other retailers the tools to enhance their inventory planning capabilities. And now, Impact Analytics just closed a strategic deal with PVH (parent company of luxury brands Tommy Hilfiger, Calvin Klein, True & Co) to build out AI solutions for assortment planning and pricing optimization. Impact Analytics' cutting edge AI and ML guided forecasting engine is built entirely on Google Cloud! Read more

Podimetrics raises $45M Series C round
Podimetrics, creator of the FDA-cleared SmartMat and integrated clinical care services team, is dedicated to early detection and prevention of diabetic amputations, one of the most debilitating and costly complications of diabetes. Its clinical care services platform leverages Google Cloud services to engage with patients, by helping save limbs, lives, and money - all while keeping vulnerable populations healthy in their own homes. Read more about their Series C funding round.

Anvyl raises +$15M in an oversubscribed Series B funding round
Congrats to Anvyl for raising a hugely successful Series B as they modernize & transform the supply chain technology market and more than doubled revenue in the last year. Read more.

Helios has kicked off 2022 in a big way
The audio tone analysis platform, Comprehend: Elite, that they provide to Wall Street quantitative hedge funds now covers all US equities and is fully available here. It’s entirely powered by Google Cloud!

Dapper Labs uses Google Cloud for performance, reliability and decentralization
In case you missed it before the holidays, Dapper Labs is working with Google Cloud as its hyperscale cloud partner to ensure performance, reliability and decentralization for the next wave of mainstream users on Flow, without needing to compromise on decentralization or sustainability. Find out more.

Geotab’s Intelligent Transportation Systems (Geotab ITS) is built on Google Cloud.
Geotab uses GKE, BigQuery, Dataflow and Cloud Composer to build an innovative solution combining analytics and access to massive data volumes so municipalities can make better transportation planning decisions. The sheer volume of information that it handles, along with a need for highly scalable and flexible tools to manage, store, and analyze that data, led Geotab to invest in Google Cloud technology. Read more.

Mux CEO shares advice for getting started with video
Mux CEO, Jon Dahl, sat down with Google Cloud Director, Nirav Sheth, to share best practices and strategies for getting started with video, along with insights and advice from his learnings as a startup founder. Listen to what he has to say.

Google Cloud is proud to support Unstoppable Women of Web3
Unstoppable Women of Web3 (UWOW3) is an action oriented community made of industry leaders supporting education & opportunities for girls, women, and minorities in this burgeoning industry. This International Women’s Day, March 8th, you can catch live interviews with Tech and Web3 leaders from all over the world, covering topics such as how to build communities, how to learn more about Web3, developing technology on the blockchain, how to talk about complex ideas with kids, and more! How you can engage:

Puppet CTO increases development speed
Hear Puppet CTO Deepak Giridharagopal discuss how they managed to build Puppet's first Saas product, Relay, fast while also ensuring they would be able to remain agile if growth was to happen quickly. Watch video.

Vimeo builds a fully responsive video platform on Google Cloud
The video platform @Vimeo leverages managed database services from Google Cloud to serve up billions of views around the world each day. Read how it uses Cloud Spanner to deliver a consistent and reliable experience to its users no matter where they are. Find out more. 

Nylas improved price-performance by 40%
You don't have to choose between price-performance and x86 compatibility. Hear from David Ting, SVP of Engineering and CISO at @nylas, to learn how Google's x86-based Tau VMs delivered 40% better price-performance than competing Arm-based VMs. Watch now.

Optimizely partners with Google Cloud on experimentation solutions 
Build the next big thing with @Optimizely Experimentation on Google Cloud - driving innovation and next-gen experimentation for enterprise companies and marketers. Check it out.

Hands-on learning lab: Stream Google Cloud data into Splunk Cloud

11 avril 2022 à 18:00

Splunk and Google Cloud customers, this one’s for you: The first Hands-on-Lab of Splunk on Google Cloud is now live and ready for enrollees. 

If you haven’t tried it yet, Google Cloud Skills Boost provides hands-on educational experiences so you can learn what you need to know about operating in the cloud. Labs from Google Cloud Skills Boost give users more than just a sandbox environment — they offer live Google Cloud projects for truly interactive learning. Users get to pick experiences ranging from short, 30-minute labs all the way up to multi-day quests to help them tailor learning to their specific needs. 

Splunk offerings on Google Cloud Platform (GCP) provide rich capabilities that cover a broad set of security scenarios, including end-to-end visibility across cloud, on-premises, and hybrid environments. Using Splunk on GCP, you can gain real-time visibility across Google Cloud events, logs, performance metrics, and billing data. Splunk also enables fast security investigations, alerting, and deeper forensic analysis to accelerate incident resolution. You can better build your security infrastructure using Splunk Phantom Apps for Google Vault, Google Workspace, Google Workspace for Gmail, and Safe Browsing. 

Now, the “Getting Started with Splunk Cloud Getting Data In (GDI) on Google Cloud” hands-on-lab is available to take you through core scenarios for data ingestion and data input in Google Cloud, enabling you to get practical, hands-on experience for all scenarios in just 90 minutes or less.

With this hands-on-lab, you’ll learn how to get streaming data from your Google Cloud environment into Splunk Cloud so your organization can leverage Splunk’s Data-to-Everything platform. The lab guides users through the installation of key Splunk components that enable you to stream data into Splunk Cloud platform:

The lab also guides you through managing the following Google Cloud resources:

As you begin the lab, you’ll launch a Dataflow job using the Splunk-specific template, configure the data inputs in Technical Add-on for Google Cloud Platform, perform sample Splunk searches across ingested data, and monitor and troubleshoot Dataflow pipelines. This enables Splunk admins to collect, analyze, and extract insights from all of your Google Cloud data in an easy-to-use and powerful way. 

Below is an architecture diagram showing the principal components and the API relationship used in the lab. In addition to Dataflow-based ingestion for Splunk, you’ll practice with Pub/Sub and K8s connector, as well as pulling data using Splunk Add-on for GCP.

This hands-on-lab provides a full-stack practice experience with Splunk on Google Cloud as part of data ingestion and processing. If you’re interested in getting started, please follow the guide here:

Getting Started with Splunk Cloud GDI on Google Cloud  

Looking Ahead with GCP and Splunk 

Stay tuned for the next Google Cloud and Splunk hands-on lab announcement, and in the meantime, check out our official Getting Data In (GDI) guide to learn about the integration after completing the lab. To take a step further and learn more about automating the process, take a look at our export Terraform module with Splunk.

How Box is unlocking multimodal enterprise agents with Gemini Embeddings 2

18 août 2026 à 18:00

Enterprise content management is experiencing its biggest architectural shift since the cloud migration era. 

For years, enterprises have stored trillions of gigabytes of critical data in Box: financial models, clinical trial protocols, M&A due diligence rooms, engineering schematics, and legal compliance playbooks. Up to this point, text-based search and retrieval-augmented generation (RAG) have successfully unlocked the vast narrative knowledge within these repositories, establishing a powerful and highly effective baseline for enterprise AI intelligence.

Traditional RAG architectures have mastered text processing, but the agentic era demands more. The next logical evolution is to extend this framework to capture the inherently multimodal, deeply spatial, and highly structured elements that exist alongside text. While text embeddings excel at indexing prose, multimodal architectures unlock a major new capability: For example, they preserve the strict row-column semantics of financial tables, interpret visual evidence like clinical data, and map the logic of multi-page flowcharts without losing their spatial layout.

To deliver next-generation capabilities that can handle the vast universe of digital content, Google Cloud and Box are integrating advanced multimodal capabilities into Box's Agentic Platform, powered by Gemini Multimodal Embeddings 2 merging Box’s industry-leading Intelligent Content Management platform with Google Cloud’s advanced AI embeddings.

Benefits of improved embedding: Extending the dimensions of document content

  1. Preserving visual and spatial geometry: Complex document elements like multi-column tables or financial matrices rely on their spatial layout to convey meaning. Converting these elements into a flat string of text can disassociate column headers from their corresponding data points. Multimodal embeddings allow systems to interpret the document exactly as a human does, maintaining the integrity of spatial relationships.

  2. Illuminating the visual modality: Enterprise documents are filled with visual indicators: technical charts, process flowcharts, branding assets, and product photography. Multimodal capabilities ensure that these elements are no longer invisible to search systems, allowing users to query images and text simultaneously.

  3. Connecting hybrid file formats: Real-world business workflows rarely live in a single document format. An agent may need to cross-reference a PDF policy, a spreadsheet tracking log, and a presentation deck. Extending RAG with multimodal embeddings creates a unified understanding across these varied formats.

The Architectural Solution: Gemini Multimodal Embeddings 2

Google Cloud’s Gemini Multimodal Embeddings 2 introduces a unified, multimodal vector space capable of embedding text, raster images, document pages, rendered spreadsheet tables, and visual charts into the same semantic representation space.

GIF_1

Key product capabilities unlocked by gemini-embeddings-2:

  • Crossmodal retrieval (text-to-visual / visual-to-text): Enables natural language queries to retrieve highly specific visual components, such as locating a target chart or diagram within a massive library of slides, without requiring manual tagging.

  • Layout-aware document embedding: Rather than breaking files into arbitrary text blocks, the system can embed document page renderings directly, preserving visual hierarchies, callout boxes, and structural context.

  • Heterogeneous format bridging: Native support for seamlessly bridging content across .docx, .xlsx, .pdf, .pptx, .png, and .csv without losing modality-specific structural information.

Three core patterns of multimodal enterprise agents

By leveraging multimodal embeddings within Box, we have identified three uniqueprimary design patterns that illustrate how organizations can extend traditional RAG to support complex, visual workflows.

Pattern 1: Complex financial & analytical reporting

The challenge

Corporate finance, research, and audit teams analyze highly structured documents where vital data resides in embedded tables, growth charts, and footnote annotations. Text-only indexing can separate these numbers from their context, making automated analysis challenging.

The multimodal advantage

  • Structural alignment: The embedding model captures the physical structure of tables and charts, allowing financial agents to understand that a column header applies to a specific row of metrics.

  • Visual trend analysis: Agents can cross-reference written summaries with visual trends in accompanying bar or line charts, identifying and pointing out discrepancies between written claims and source data.

  • Contextual sourcing: Users can query complex portfolios and instantly retrieve the exact page, table, or chart supporting a specific metric.

2

Pattern 2: Multimodal clinical decision support & assisted diagnosis

The challenge

In healthcare and clinical environments, critical patient data is fragmented across vastly different, unstructured visual and textual formats — ranging from external physical photos (visual evidence) and microscopic pathology slides (lab reports) to structured risk matrices (triage grids). Traditional text-based systems or isolated analysis tools cannot synthesize these cross-modal relationships simultaneously, which can delay critical diagnoses or risk missing immediate, life-threatening procedural complications.

The multimodal advantage

  • Cross-modal clinical synthesis: Evaluates physical symptoms alongside cellular-level laboratory evidence simultaneously by indexing clinical photos, histopathology imagery, and triage grids into a single space.

  • Granular anomaly identification: Connects niche visual patterns under a microscope (like parasitic cyst walls) with medical knowledge to rapidly isolate rare conditions.

  • Risk-aware decision support: Cross-references findings against triage frameworks to deliver instant warnings about immediate patient risks, such as life-threatening anaphylactic shock.

3

Pattern 3: Cross-document multimodal synthesis & data reconciliation

The challenge

Enterprise information is fragmented across disconnected files and formats (e.g., PDF minutes, Excel charts, PNG flyers, and email threads). Traditional tools analyze these files in isolation, failing to connect the dots when verifying details or resolving data contradictions across independent documents.

The multimodal advantage

  • Cross-file synthesis: Connects information across entirely different formats (PDFs, spreadsheets, images, emails) simultaneously to answer complex business queries.

  • Conflict resolution: Flags and resolves contradictions between assets, such as catching outdated pricing on an image by cross-checking it against the latest financial spreadsheets.

  • Visual-to-text auditing: Audits visual or scanned files against text-based records (e.g., verifying a signed PDF contract against a legal review email) to catch missing clauses or changes.

4

The future of agentic enterprise content management

The integration of gemini-embeddings-2 into Box’s Agentic Platform is an important new capability to improve the next era of content intelligence. Multimodal embeddings help Box to move beyond basic search to active, intelligent collaboration.Box's Intelligent Content Management platform represents a fundamental shift in enterprise AI infrastructure — moving beyond passive document storage to deliver a governed, semantically indexed reasoning layer where AI agents can interrogate, cross-reference, and act on content with full compliance and security controls already in place. 

Powered by multimodal embeddings and a suite of native AI agents spanning search, metadata extraction, research, analysis, and composition, Box enables organizations to proactively surface insights such as flagging stale pricing data, expiring contract clauses, or cross-document contradictions before they become business risks. For high-complexity industries like financial services, life sciences, and legal operations, Box's ability to reason across text, tables, charts, and images makes multimodal understanding a competitive requirement. 

Designed to interoperate with the broader enterprise AI ecosystem, Box serves as the single governed content foundation that ensures every AI-driven workflow is grounded in authorized, auditable enterprise data.

When you think about it, the enterprise data landscape was always multimodal. Now we have the technology to make the most of it. By integrating gemini-embeddings-2, Box helps its users unlock unprecedented value from unstructured enterprise content. Product leaders who embrace multimodal-first architectures, rigorous precision benchmarking, and audit-ready grounding will lead the next wave of enterprise productivity and innovation.

The team would like to thank Ken Ikeda, Afshaan Mazagonwalla, and Samip Thakkar for their work on this project.

Accelerating automotive innovation with C4A-metal and Panasonic Automotive vSkipGen

20 juillet 2026 à 18:00

As the automotive landscape accelerates toward software-defined vehicles, Cockpit Domain Controllers (CDCs) are becoming the core of next-generation in-cabin experiences. The ability to rapidly develop, test, and validate CDC software in a flexible, hardware-independent environment is critical for innovation and time-to-market. However, physical hardware constraints and the requirement for high-performance graphics present significant challenges for global development teams. 

Panasonic Automotive’s vSkipGen™ addresses these challenges as a next-generation CDC virtualization platform, now validated on Google Cloud’s C4A-metal, our Axion bare-metal offering. By integrating Panasonic Automotive’s advanced Unified HMI™ remote GPU offload technology with support for Android Automotive OS (AAOS) and Android SDV, vSkipGen delivers a robust, cloud-native solution for cockpit software development and validation — empowering teams to innovate without hardware limitations.

At Google Cloud, we provide workload-optimized infrastructure to help ensure the right resources for every task. Similar to the entire Axion virtual machine family, C4A-metal instances are built on Google Cloud’s custom Arm-based Axion architecture. C4A-metal offers 96 vCPUs, two DDR5 memory configurations (384GB and 768GB), and up to 100Gbps of networking bandwidth. It also provides full support for Google Cloud Hyperdisk, including Balanced, Extreme, Throughput, and ML types. And like the rest of the bare metal portfolio, C4A-metal is powered by Titanium, a key component for multi-tier offloads and security that is foundational to our infrastructure. 

High performance for demanding workloads

C4A-metal is particularly well-suited for complex tasks such as creating digital twins of vehicle cockpits where performance must accurately mirror real-world behavior. Traditionally, the transition to software-defined vehicles has relied on expensive and scarce physical prototypes; C4A-metal overcomes this by offering the high performance and hardware-level access of bare metal with the scalability of the cloud. 

Panasonic Automotive leverages C4A-metal to bypass traditional hardware bottlenecks, enabling their teams to execute complex virtualization tasks and accelerate the development of next-generation cockpit software.

"Google Cloud’s Axion Bare Metal has been a game-changer for our vSkipGen™ platform. By providing scalable, high-performance Arm-based infrastructure, C4A-metal allows our teams to develop and test production-intent software in the cloud with behavior that closely matches target automotive hardware. This cloud-to-car bit parity reduces dependence on costly physical prototypes, improves validation efficiency, increases test coverage and accelerates time-to-market for next-generation cockpit platforms.” - Andrew Poliak, CTO, Panasonic Automotive Systems America.

By leveraging vSkipGen and Unified HMI on C4A-metal, automotive manufacturers can now build, test, and validate full AAOS stacks in a hardware-independent, cloud-native environment, moving from physical dependency to scalable digital twins.

1

Figure 1: Unified HMI solution overview

How vSkipGen works: Virtualizing the cockpit with Cuttlefish

Panasonic Automotive’s vSkipGen acts as a digital twin for physical CDC hardware. To provide a hardware-agnostic environment for Android virtual machines, vSkipGen uses components of Android Cuttlefish. At its core, vSkipGen leverages a cloud-optimized Virtual Machine Monitor (VMM) built on crosvm (the open-source, security-focused VMM originally developed for Chrome OS) which utilizes Linux KVM (Kernel-based Virtual Machine) for hardware-assisted virtualization. The VMM backend is implemented in Rust for enhanced security, scalability, and performance. By running the full stack on C4A-metal (see Figure 2), Panasonic lets developers boot a full AAOS image in the cloud, which behaves exactly like the software running in a physical vehicle.

The platform virtualizes all essential peripherals, such as the audio, GPU, sensors, cameras, Controller Area Network (CAN), Bluetooth, and Wi-Fi, using the VirtIO standard. This VirtIO-native approach allows developers to interact with the virtual devices exactly as they would with the physical hardware. Furthermore, vSkipGen seamlessly connects with automotive simulators and software-in-the-loop (SiL) environments for comprehensive scenario and edge-case validation, enabling teams to conduct software validation and run automated test suites without needing early access to physical prototypes.

2

Figure 2: vSkipGen Cockpit virtualization architecture

3

Figure 3: Remote GPU rendering flow

Accelerating graphics on the go with Unified HMI

High-performance graphics is central to the modern driving experience, but rendering it in a virtual environment can be challenging. Panasonic’s Unified HMI solves this by decoupling HMI rendering from specific hardware. A lightweight Unified HMI component operates outside the VM to offload OpenGL ES commands (the specific data being rendered) from the Cuttlefish instance to GPU-equipped compute resources on Google Cloud, which handle the workloads with hardware acceleration. The rendered UI is then streamed to any standard browser using low-latency WebRTC. This helps ensure that global development teams can experience high-fidelity visuals in real time, regardless of their location.

Unified HMI establishes a unified virtual display layer across multiple Electronic Control Units (ECUs) and virtual machines, allowing applications to render to different displays from anywhere within the system.

Benefits for software-defined vehicle development

With C4A-metal and Panasonic Automotive’s vSkipGen, manufacturers building software-defined vehicles can: 

  • Build and validate AAOS-based software in the cloud using Cuttlefish before physical hardware is available

  • Use industry-standard VirtIO to emulate critical CDC devices for robust, production-grade validation

  • Stream interactive cockpit experiences to any browser to support distributed engineering teams

  • Run multiple isolated CDC instances in parallel to support large-scale automated testing and CI/CD pipelines

  • Reduce cost and environmental impact by minimizing the need for expensive physical hardware prototypes, supporting sustainable development practices

  • Enjoy a future-ready architecture built on open standards, crosvm, and Rust for enhanced security, performance, and long-term adaptability

Get started

C4A-metal is generally available worldwide; please refer to the public documentation for additional information. Panasonic Automotive’s vSkipGen with Unified HMI for Google Cloud will soon be available for evaluation access. To explore how these solutions can help you speed up cockpit software development, contact the team at vSkipGenSupport@panasonicautomotive.com. 


Disclaimer
All trademarks, tradenames, and service marks used herein are the property of their respective owners.

Claude at scale on Google Cloud: Frontier AI, built for enterprise production

14 juillet 2026 à 18:00

Running frontier AI in production is demanding — accelerators to manage, latency to hold steady across continents, regulated data to keep in-region, and long-context requests to serve reliably. Claude on Google Cloud is built for exactly this. 

Like Monet and water lilies, frontier models and the enterprise platforms are often better together. In our case, Claude brings the reasoning, and Google Cloud brings the managed infrastructure, global reach, and compliance posture that enterprises already run on. Calling Claude becomes operationally identical to calling any other Google Cloud service — same Identity and Access Management (IAM), same VPC Service controls, same observability — so teams are able to spend their time building features instead of running inference infrastructure.

This post walks through what Claude on Google Cloud delivers in production across four areas: 

  1. Managed infrastructure that gives engineers their time back 

  2. Global endpoints that hold latency low, and uptime high for a worldwide user base 

  3. Security and data-sovereignty controls inherited straight from Google Cloud

  4. Serving-layer features that keep cost and performance optimized at scale.

Managed infrastructure that frees engineering time

Claude on Google Cloud runs on fully managed infrastructure, so enterprise teams ship features instead of building clusters. Compute provisioning, auto-scaling logic, load balancing, and failover at frontier-model scale are handled by the platform — work that would otherwise occupy multiple teams full-time. 

Claude is available through Agent Platform's Model Garden as a Model-as-a-Service offering, ready to use over standard REST / JSON over HTTP/1.1 or HTTP/2 endpoints. Invoking Claude is operationally identical to invoking any other Google Cloud service: the same IAM policies, the same VPC controls, and the same observability stack via Cloud Logging and Cloud Monitoring. 

Serving Claude takes a few lines of Python using the AnthropicVertex client:

code_block
<ListValue: [StructValue([('code', 'from anthropic import AnthropicVertex\r\n\r\nclient = AnthropicVertex(\r\n project_id="your-project-id",\r\n region="us"\r\n)\r\n\r\nmessage = client.messages.create(\r\n model="claude-opus-4-8",\r\n max_tokens=1024,\r\n messages=[{"role": "user", "content": "Analyze this system architecture."}]\r\n)'), ('language', ''), ('caption', <wagtail.rich_text.RichText object at 0x7fd94a6a1d30>)])]>

The same AnthropicVertex client handles prompt caching, tool use, structured outputs, streaming, and adaptive thinking; for batch inference, use Vertex AI Batch Prediction. Authentication uses Application Default Credentials; requests automatically inherit your project's IAM and VPC configuration. 

Global reach with consistent latency and built-in failover

Serving a worldwide user base from a single endpoint produces high tail latency and a single point of failure. Most enterprises can't replicate inference infrastructure across continents while keeping performance consistent.

Agent Platform exposes three endpoint types for Claude, each solving a different production requirement:

  1. Global endpoints route requests to a region with available AI compute capacity. For example, if us-central1 is capacity-constrained, traffic redirects to europe-west1 or another region with available capacity. That’s automatic failover and geographic load balancing without application-side routing logic. Global endpoints are ideal for maximum availability and lowest cost.

  2. Regional endpoints like us-east5 or europe-west1 keep prompts, completions, and intermediate state inside a specific geographical boundary, making it ideal for low latency and data-residency requirements.

  3. Multi-region endpoints give U.S. or EU data residency without single-region dependency. They dynamically route across regional endpoints  providing built-in resilience against regional outages and capacity constraints.

The diagram below shows how applications reach Claude through these endpoint types, and how the Agent Platform serving layer routes traffic to the Compute AI clusters across regions:

1 _GC_BlogGraphics_Anthropic

Serving Claude Models From Regional & Global Endpoints

2_GC_BlogGraphics_Anthropic

Serving Claude Models From Multi-Region Endpoints

3_GC_BlogGraphics_Anthropic

Serving Claude Models From Regional Endpoints

Enterprise security and data sovereignty built in

Regulated workloads — financial services, healthcare, and government — get enterprise-grade security and data sovereignty without trading compliance for convenience, and without re-engineering the hardest layer to control: inference, where prompts, completions, and intermediate state all flow through the serving stack.

Claude on Agent Platform inherits Google Cloud's full security posture. FedRAMP High and HIPAA compliance enable deployment in government, healthcare, and financial services environments. VPC Service Controls let organizations define a perimeter around Agent Platform resources, preventing data exfiltration. IAM-native access control governs Claude endpoints with the same roles and policies that protect every other Google Cloud resource — no separate API keys to manage or rotate. Cloud Logging and Cloud Monitoring provide near real-time visibility into token usage, error rates, latency, and quota consumption.

Combined with the regional and multi-region endpoints above, this gives regulated customers a path to running frontier AI in production without re-auditing their compliance posture.

Optimized for cost and performance at scale

In production, cost and performance drive every architectural decision. Getting both right requires capabilities from two layers: Claude's native model features, and Google Cloud's serving infrastructure. Agent Platform supports both, so teams can optimize across the stack without managing them separately.

Claude-native capabilities, fully supported on Agent Platform

These features are built into Claude and available on Agent Platform without any additional configuration:

  • Prompt caching stores and reuses shared prefixes — long system prompts, legal documents, codebases — reducing request latency by up to 80% and cost by up to 90%.

  • Streaming responses over server-sent events deliver tokens as they are generated, critical for chat interfaces and coding assistants where perceived latency matters.

  • Extended and adaptive thinking lets Claude dynamically determine when and how much to reason through complex, multi-step problems — and allows users to dial the thinking effort directly, for example to control cost. Optimized for use cases like advanced code generation, mathematical reasoning, and multi-document analysis.

  • Extended context windows up to 1M tokens (for Claude Opus 4.6,Sonnet 4.6 and newer models) enable long-document analysis, large codebase reasoning, and multi-turn conversations at depth.

Google Cloud serving infrastructure

Agent Platform adds its own serving-layer capabilities on top of Claude's native features:

  • Batch prediction handles large-scale offline workloads — document classification, content moderation, bulk summarization — asynchronously at lower priority and reduced cost.

  • Provisioned throughput reserves dedicated inference capacity for mission-critical workloads, isolating them from public traffic and ensuring predictable performance during peak demand.

  • Memory management and scheduling for long-context requests is handled at the infrastructure layer,.

Together, these two layers give teams the full range of optimization levers — from model-level efficiency to infrastructure-level capacity control — on a single, unified platform.

From inference to agents

The same infrastructure that serves Claude inference powers the agent layer of Agent Platform on Google Cloud. The build-and-register flow has three steps:

  1. Build with Claude. Claude is well-suited as an orchestration backbone — its extended context window, native tool use, and adaptive thinking make it effective at planning multi-step tasks and delegating to sub-agents. Pick Claude Opus, Sonnet, or Haiku from the Model Garden, then build with the Agent Development Kit (ADK) — code-first in Python, Go, Java, or TypeScript — deploy to Agent Runtime, Cloud Run or Google Kubernetes Engine.

  2. Deploy the Agent to a Runtime. Depending on your use case, select Agent Runtime, Google Kubernetes Engine or GKE Agent Sandbox to run your deployed agents.

  3. Interoperate over A2A. The Agent2Agent protocol runs at 150+ organizations, letting a registered Claude-powered agent delegate tasks to agents from SaaS and other service providers.

The result: a planning agent built on Claude can orchestrate sub-tasks across the broader agent ecosystem, under unified IAM, fully auditable, on the same infrastructure that serves the underlying inference.

Start building

Open the Agent Platform console, enable Claude in the Model Garden, and make your first API call with the AnthropicVertex SDK. Add prompt caching, provisioned throughput, and other features as your workload demands. When you're ready to go agentic, learn more about Claude on Agent Platform.

Reach out to your Google Cloud sales representative to discuss bringing Claude into your production environment at scale.

❌