<?xml version="1.0" encoding="utf-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:media="http://search.yahoo.com/mrss/"><channel><title>AI &amp; Machine Learning</title><link>https://cloud.google.com/blog/products/ai-machine-learning/</link><description>AI &amp; Machine Learning</description><atom:link href="https://cloudblog.withgoogle.com/blog/products/ai-machine-learning/rss/" rel="self"></atom:link><language>en</language><lastBuildDate>Mon, 31 Aug 2026 15:53:00 +0000</lastBuildDate><image><url>https://cloud.google.com/blog/products/ai-machine-learning/static/blog/images/google.a51985becaa6.png</url><title>AI &amp; Machine Learning</title><link>https://cloud.google.com/blog/products/ai-machine-learning/</link></image><item><title>What’s new in AI infrastructure and orchestration in August</title><link>https://cloud.google.com/blog/topics/ai-infrastructure/whats-new-in-ai-infrastructure-this-month/</link><description>&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Welcome back to &lt;/span&gt;&lt;span style="font-style: italic; vertical-align: baseline;"&gt;What’s new in AI infrastructure and orchestration this month&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;, a collection of product updates, how-tos, customer stories, research and other resources about all the AI compute, networks, storage, frameworks, and orchestration software that you can find at Google Cloud. To be honest, we thought August would be a slow month, but nothing could be further from the truth. Read on and you’ll see what we mean.&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;August 2026&lt;/span&gt;&lt;/h3&gt;
&lt;h4&gt;&lt;span style="vertical-align: baseline;"&gt;Product, technology, and tools updates&lt;/span&gt;&lt;/h4&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Product update:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; &lt;/span&gt;&lt;a href="https://cloud.google.com/filestore"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Filestore&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;, Google Cloud’s first-party, secure, scalable NFS file service, has emerged as a popular storage platform for AI and agentic workflows, and now, it’s even better suited to the task, with a new backend storage layer built directly on &lt;/span&gt;&lt;a href="https://cloud.google.com/blog/products/storage-data-transfer/how-colossus-optimizes-data-placement-for-performance?e=48754805"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Colossus&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;, Google’s foundational distributed storage system. This new backend lets you provision IOPS independently from storage capacity, and is deeply integrated with GKE. In AI environments, this can help you service so-called agentic swarms — large groups of agents that need to read and write to a common dataset — without a drop off in performance. For more, check out the &lt;/span&gt;&lt;a href="https://cloud.google.com/blog/products/storage-data-transfer/filestore-file-service-runs-on-colossus?e=48754805"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;blog post&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;. &lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;New feature: &lt;/strong&gt;&lt;a href="https://cloud.google.com/blog/products/containers-kubernetes/gvisor-sandboxes-for-ray-clusters-on-gke?e=48754805"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;gVisor sandboxes are now available in distributed Ray clusters on GKE&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;. In partnership with Anyscale, we introduced an experimental library for Ray that brings gVisor, Google’s open-source application kernel, directly into distributed Ray clusters. gVisor provides lightweight environments with stronger isolation than ordinary containers, plus fast startup times and low memory overhead. To try out these sandboxing capabilities on GKE, head over to the &lt;/span&gt;&lt;a href="https://docs.ray.io/en/master/cluster/kubernetes/examples/ray-sandboxing.html" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Ray sandboxing User Guide&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Product update: &lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;Looking for high-performance, easy-to-use infrastructure on which to run a personal AI agent, but don’t want to spend a lot of money? New &lt;/span&gt;&lt;a href="https://cloud.google.com/blog/products/serverless/introducing-cloud-run-instances"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Cloud Run instances&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; are dedicated, singleton compute runtimes on Cloud Run that won’t shut down when the agent is idle. Better yet, the cost to run a Cloud Run instance with 1 vCPU and 1 GiB of memory continuously for 30 days is just $5.70.  &lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h4&gt;&lt;span style="vertical-align: baseline;"&gt;Practitioner guides, documentation and how-tos&lt;/span&gt;&lt;/h4&gt;
&lt;ul&gt;
&lt;li role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;How-to guide:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Big news in Model Context Protocol (MCP) land: As of the 2026-07-28 specification, the protocol core is “completely stateless. The handshake is gone. The initialize / initialized handshake (SEP-2575) and the logical Mcp-Session-Id header (SEP-2567) have been removed entirely. Instead, every request is now self-describing and independent.” Whoa. Learn more about the changes that the latest MCP specification brings, and more importantly, how to implement them, in &lt;/span&gt;&lt;a href="https://developers.googleblog.com/scaling-ai-agent-infrastructure-with-the-mcp-stateless-updates/" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;this Google Developers blog&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;.  &lt;/span&gt;&lt;/li&gt;
&lt;li role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Guide: &lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;Real-time AI systems make a mess of traditional network load balancing techniques.&lt;/span&gt;&lt;span style="font-style: italic; vertical-align: baseline;"&gt; “Instead of handling isolated requests, the backend has to manage a continuous, live bidirectional stream. You’re dealing with a constant stream of audio chunks, transcripts, model outputs, and synthesized speech flowing back and forth simultaneously.”&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt; Things only get worse when the user gets involved. &lt;/span&gt;&lt;span style="font-style: italic; vertical-align: baseline;"&gt;“The server has to immediately halt its current speech generation, pivot to update the context, maybe trigger a new tool, and start drafting a different response; this must be done without dropping the connection.”&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt; For a new approach to managing load in the AI era, read &lt;/span&gt;&lt;a href="https://developers.googleblog.com/scaling-real-time-ai-agents-with-session-aware-load-balancing/" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Scaling real-time AI agents with session-aware load balancing&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;.&lt;/span&gt;&lt;/li&gt;
&lt;li role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;How-to:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Learn how to build an elastic, scalable LLM inference platform on GKE, even with a mix of different GPU accelerators. The proposed architecture combines Capacity Advisor and Compute Advisor, plus high-performance storage like RunAI:model streamer or GCPFuse with parallel downloads. Get all the details &lt;/span&gt;&lt;a href="https://discuss.google.dev/t/how-to-build-an-elastic-scalable-llm-inference-platform-on-gke-using-fluid-compute/388108" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;here&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;.&lt;/span&gt;&lt;/li&gt;
&lt;li role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Documentation: &lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;The thing about hosts with GPUs or TPUs is that you can’t use live migration to update them, setting up a maintenance challenge. In this new docs page, learn how to &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/kubernetes-engine/docs/how-to/perform-host-maintenance-accelerators"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;update accelerator-equipped hosts&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; according to your tolerance for downtime for your training and inference workloads.   &lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt; &lt;/strong&gt;&lt;/li&gt;
&lt;li role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Documentation: &lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;Advanced Compute Images, or ACIs, are standardized image stacks for AI/ML and HPC infrastructure, so you don’t need to manually build your own custom images. In this new docs page, learn how to &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/compute/docs/instances/use-aci-images"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;create an ACI image&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; using the Google Cloud CLI, console, or SchedMD's Slurm workload manager&lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;. &lt;/strong&gt;&lt;/li&gt;
&lt;li role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Guide: &lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;AI workloads are notoriously difficult to architect, resource-intensive, and bursty, which can also lead to scaling bottlenecks and large pools of underutilized — or misutilized — compute resources. A new blog outlines the &lt;/span&gt;&lt;a href="https://cloud.google.com/blog/topics/ai-infrastructure/best-practices-for-dynamic-capacity-management?e=48754805"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;three main ways to achieve dynamic capacity management in Google Cloud&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;: 1) scheduling capacity for planned downtime; 2) maintaining automated fallback capacity for unplanned downtime; and 3) relying on GKE’s core orchestration capabilities to automate resource allocation. &lt;/span&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;h4&gt;&lt;span style="vertical-align: baseline;"&gt;Customer and partner updates&lt;/span&gt;&lt;/h4&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;span style="vertical-align: baseline;"&gt;Business orchestration software provider &lt;/span&gt;&lt;a href="https://www.uipath.com/" rel="noopener" target="_blank"&gt;&lt;strong style="text-decoration: underline; vertical-align: baseline;"&gt;UiPath&lt;/strong&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; was dealing with spiky workloads, and wanted more predictable costs. To get there, it re-architected its infrastructure, moving from isolated clusters to a shared Google Cloud GPU fleet that included both A3 VM instances (NVIDIA H100 GPUs) for training with G4 VM instances (NVIDIA RTX PRO 6000 Blackwell Server Edition GPUs) for inference. You can read more about their architecture &lt;/span&gt;&lt;a href="https://cloud.google.com/blog/topics/customers/how-uipath-built-its-high-performance-gpu-platform"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;here&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;. &lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;a href="https://mirendil.com/" rel="noopener" target="_blank"&gt;&lt;strong style="text-decoration: underline; vertical-align: baseline;"&gt;Mirendil&lt;/strong&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;, an frontier AI lab focused on accelerating AI development, announced that it is &lt;/span&gt;&lt;a href="https://cloud.google.com/blog/topics/startups/mirendil-selects-ai-hypercomputer?e=48754805"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;using AI Hypercomputer&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; with both TPUs and NVIDIA GPUs to support its model pre-training and post-training applications. &lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;a href="https://replen.it/" rel="noopener" target="_blank"&gt;&lt;strong style="text-decoration: underline; vertical-align: baseline;"&gt;Replenit&lt;/strong&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;, a retail CRM provider, built its AI decision engine in Google Cloud, using BigQuery, Gemini Enterprise Agent Platform, and open-source Gemma models that it runs on Cloud TPUs. This latter combination provided Replenit with 90% lower pipeline costs than their previous cloud provider, the company reports. Read the &lt;/span&gt;&lt;a href="https://cloud.google.com/customers/replenit?e=48754805&amp;amp;hl=en"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;full case study&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; for more. &lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;a href="https://www.malachyte.com/" rel="noopener" target="_blank"&gt;&lt;strong style="text-decoration: underline; vertical-align: baseline;"&gt;Malachyte&lt;/strong&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; architected its AI-powered e-commerce recommendation platform on top of Bigtable, Managed Service for Apache Kafka, Pub/Sub, Compute Engine, and last but not least, GKE. See how it all comes together in &lt;/span&gt;&lt;a href="https://cloud.google.com/blog/products/data-analytics/solving-retails-cold-start-problem-malachytes-recommendation-reinvention?e=48754805"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;this blog&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;hr/&gt;
&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;July 2026&lt;/span&gt;&lt;/h3&gt;
&lt;h4&gt;&lt;span style="vertical-align: baseline;"&gt;Product, technology, and tools updates&lt;/span&gt;&lt;/h4&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Product update:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; &lt;/span&gt;&lt;a href="https://cloud.google.com/products/managed-lustre"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Google Cloud Managed Lustre&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; is now GA, and available in four distinct performance tiers that deliver throughput ranging from 125 MB/s, 250 MB/s, 500 MB/s, to 1000 MB/s per TiB of capacity — with the ability to scale up to 8 PB of storage capacity. The Managed Lustre solution is powered by DDN’s EXAScaler, combining DDN's decades of leadership in high-performance storage with Google Cloud's expertise in cloud infrastructure.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Product update:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; &lt;/span&gt;&lt;a href="https://cloud.google.com/blog/products/compute/c4n-network-and-storage-optimized-vms?e=0"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;C4N network and storage optimized VMs are now GA&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;. C4N is our first network- and block-storage-optimized VM series built to eliminate data-transfer bottlenecks. Powered by 5th Gen Intel Xeon Scalable processors and built on Google's &lt;/span&gt;&lt;a href="https://cloud.google.com/titanium?e=0"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Titanium&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; offloading hardware, it achieves 400 Gbps network bandwidth, 95 million packets per second (MPPS), and up to 25 GiB/s of block storage throughput when paired with Hyperdisk Extreme.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;New feature:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/kubernetes-engine/docs/concepts/planning-large-clusters#clusters-5k-nodes"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;GKE Dataplane V2 up to 15K Nodes with Network Policies (GA)&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;. This capability enables standard GKE clusters to scale up to 15,000 nodes while maintaining full active Network Policy enforcement, supporting the massive infrastructure needs of large enterprise and AI/ML customers.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;New feature:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; &lt;/span&gt;&lt;a href="https://cloud.google.com/blog/products/containers-kubernetes/introducing-co-operative-time-slicing-for-rl-in-llm-d?e=0"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Co-operative time-slicing in llm-d&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;. If you’re running reinforcement learning (RL) workloads, you can now interleave independent RL jobs onto shared physical hardware, increasing aggregate accelerator duty cycles from a ~40% baseline up to 70% without impacting model convergence or accuracy. &lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;New AI security tool:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; &lt;/span&gt;&lt;a href="https://cloud.google.com/blog/products/identity-security/introducing-k8s-aibom-on-gke-for-automated-ai-bills-of-materials?e=0"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Looking to secure your AI supply chain on GKE&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;, deploy AI workloads safely, and cut down on shadow AI? We open-sourced k8s-aibom, a lightweight, unprivileged Kubernetes controller that continuously monitors container clusters to automatically detect running AI runtimes (like vLLM and Triton) and generate standard CycloneDX Machine Learning Bill of Materials (ML-BOMs). Check out the &lt;/span&gt;&lt;a href="https://github.com/GoogleCloudPlatform/k8s-aibom" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;k8s-aibom project&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; and get involved.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h4&gt;&lt;span style="vertical-align: baseline;"&gt;Practitioner guides and how-tos&lt;/span&gt;&lt;/h4&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;How-to guide: &lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;On July 27, Google announced &lt;/span&gt;&lt;a href="https://discuss.google.dev/t/announcing-day-0-support-for-kimi-k3-on-google-cloud/385392" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Day 0 support for Moonshot AI’s Kimi K3&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; 2.8-trillion-parameter open-weight model, the day weights were released. Whichever your preferred deployment path — via Model Garden, custom orchestration, or GKE with llm-d recipes — this guide offers detailed step-by-step instructions to help you evaluate and pilot Kimi K3 in Google Cloud. &lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;How-to guide:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; &lt;/span&gt;&lt;a href="https://cloud.google.com/blog/topics/developers-practitioners/autopilot-clusters-with-gke-managed-dranet-gpus-and-tpus"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Google Kubernetes Engine (GKE) managed DRANET supports both GPUs and TPUs&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;. There are several configurations to use this implementation, including standard cluster (where you have full control) and autopilot cluster (where Google does the heavy configs for you). Take a deeper dive in the hands-on lab, &lt;/span&gt;&lt;a href="https://codelabs.developers.google.com/codelabs/gke-autopilot-tpus-dranet-gemma#0" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;GKE Autopilot clusters with TPUs, GKE managed DRANET and Gemma 4&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;How-to guide:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Learn to run Ray on TPUs, not GPUs. In &lt;/span&gt;&lt;a href="https://developers.googleblog.com/run-ray-on-tpu-part-1-the-foundations/" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Part 1&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; of this two-part series, we discuss TPU slices (hint: Ray thinks of them as just another accelerator on which to schedule), then walk through Ray’s various AI libraries (&lt;/span&gt;&lt;a href="https://developers.googleblog.com/run-ray-on-tpu-part-2-ray-ai-libraries/" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Part 2&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;).&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;How-to guide: &lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;Evaluate TPUs for sample workloads using a new microbenchmark suite that helps you accurately assess whether a device is achieving its theoretical performance specifications, and to identify specific performance gaps or architecture-specific bottlenecks. Dive in &lt;/span&gt;&lt;a href="https://developers.googleblog.com/how-to-use-google-microbenchmarks-for-evaluating-tpu-performance/" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;here&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;. &lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;How-to guide:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Scale your agents without killing your budget. &lt;/span&gt;&lt;a href="https://cloud.google.com/blog/products/containers-kubernetes/reduce-your-agents-costs-with-gke-agent-sandbox?e=48754805"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Learn how GKE orchestration can help you safely pack more agents onto a fixed compute footprint&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; with GKE Agent Sandbox and Pod snapshots. Whether your goal is performance or cost optimization, we teach you how to turn the right dials for optimal agent efficiency. &lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Technical blueprint: &lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;Inside the optimization of Mistral 3 large inference on Ironwood. This blog outlines how one Google team optimized Mistral 3 large MoE model inference on Google’s Ironwood (TPU v7x), achieving a 1.5x performance gain. They did so with hybrid sharding, replacing linear VPU summations with tree reductions, optimizing GMM/MLA kernels, and adopting asynchronous scheduling. As a result, they boosted throughput by up to 48% while maintaining benchmark accuracy neutrality. Read the full blog &lt;/span&gt;&lt;a href="https://discuss.google.dev/t/inside-the-optimization-of-mistral-3-large-inference-on-ironwood/385847" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;here&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h4&gt;&lt;span style="vertical-align: baseline;"&gt;Research, reports and deep-dives&lt;/span&gt;&lt;/h4&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Report: &lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;Google was named a Leader in the inaugural &lt;/span&gt;&lt;a href="https://cloud.google.com/blog/topics/ai-infrastructure/google-is-a-leader-in-gartner-magic-quadrant-for-ai-infra?e=0"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Gartner&lt;/span&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;&lt;span style="vertical-align: super;"&gt;Ⓡ&lt;/span&gt;&lt;/span&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt; Magic Quadrant™ for AI Infrastructure&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;, positioned highest for ‘Ability to Execute’ and furthest for ‘Completeness of Vision’. Gartner called out Google’s proprietary scalable compute, integrated AI Hypercomputer architecture, and the scale of our AI compute capacity as key strengths. Download a copy &lt;/span&gt;&lt;a href="https://cloud.google.com/resources/content/2026-gartner-mq-ai-infrastructure?e=0"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;here&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Report:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; We recently surveyed more than 1,400 senior IT leaders for our &lt;/span&gt;&lt;a href="https://cloud.google.com/resources/content/state-of-infrastructure-in-the-agentic-ai-era?e=48754805"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;State of AI Infrastructure report&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;, and a resounding pattern emerged: The gap between AI ambition and infrastructure reality is widening. In fact, 83% of organizations say they require infrastructure upgrades to support production-grade agentic AI. &lt;/span&gt;&lt;a href="https://cloud.google.com/blog/products/compute/state-of-ai-infrastructure-report-overview?e=0"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Read the accompanying blog&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; to understand how adapting your infrastructure to meet the demands that agentic applications place on your systems will help you move from pilot to production.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;hr/&gt;
&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;June 2026&lt;/span&gt;&lt;/h3&gt;
&lt;h4&gt;&lt;span style="vertical-align: baseline;"&gt;Product, technology and tool updates&lt;/span&gt;&lt;/h4&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Product update:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Protecting sensitive data used with AI is a critical part of advanced and secure cloud infrastructure. &lt;/span&gt;&lt;a href="https://cloud.google.com/security/products/confidential-computing?e=0"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Confidential Computing&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; cryptographically protects data in use in hardware-based Trusted Execution Environments (TEEs) with verifiable data integrity, and is &lt;/span&gt;&lt;a href="https://cloud.google.com/blog/products/identity-security/verifiable-trust-in-the-ai-era-whats-new-in-confidential-computing?e=0"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;now available&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; on the accelerator-optimized &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/compute/docs/accelerator-optimized-machines#g4-series"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;G4 machine series&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;, featuring &lt;/span&gt;&lt;a href="https://www.nvidia.com/en-us/products/workstations/professional-desktop-gpus/rtx-pro-6000-family/" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;NVIDIA RTX PRO 6000 Blackwell Server Edition GPUs&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;. Get started with &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/confidential-computing/confidential-vm/docs/create-a-confidential-vm-instance-with-gpu"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Confidential G4 VMs&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; and &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/kubernetes-engine/docs/how-to/gpus-confidential-nodes"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Confidential G4 GKE Nodes&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Developer resource: &lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;The new &lt;/span&gt;&lt;a href="https://cloud.google.com/products/tpu/tpu-developer?e=0"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;TPU Developer Hub&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; is the place to go for model builders, optimizers, and developers to learn to unlock the full performance of Google Cloud TPUs. Read more in this &lt;/span&gt;&lt;a href="https://developers.googleblog.com/unlocking-the-power-of-the-tpu-stack-introducing-our-new-developer-hub/" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;blog&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;. &lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;New product: &lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;Scale your AI workloads with the new &lt;/span&gt;&lt;a href="https://discuss.google.dev/t/stop-training-blind-scaling-ai-with-the-new-opentelemetry-based-tpu-ai-telemetry-collector-agent/375210" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;OpenTelemetry-Based TPU AI Telemetry Collector Agent&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;. For the first time, you can route high-fidelity TPU hardware telemetry to Google Cloud Monitoring, Google Managed Prometheus, or your own self-hosted Grafana stack.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h4&gt;&lt;span style="vertical-align: baseline;"&gt;Practitioner guides and how-tos&lt;/span&gt;&lt;/h4&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;How-to guide:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Learn how to build high availability into an AI inference workload running on GKE Inference Gateway with TPUs, Cloud Storage FUSE and Dynamic Resource Allocation (DRA). This &lt;/span&gt;&lt;a href="https://cloud.google.com/blog/topics/developers-practitioners/experimenting-with-tpus-gke-managed-dranet-and-multi-cluster-inference-gateway?_gl=1*jj3plw*_ga*OTAxNzc0MzU1LjE3ODIyMjAxNDk.*_ga_4LYFWVHBEB*czE3ODI3NTc3NzAkbzkkZzEkdDE3ODI3NTg2MDEkajYwJGwwJGgw&amp;amp;e=0"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;blog&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; provides an overview, or you can get all the technical details in the &lt;/span&gt;&lt;a href="https://codelabs.developers.google.com/codelabs/gke-inference-gateway-multi-cluster-tpus-dranet#0" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;hands-on codelab&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;How-to guide:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Did you know you can connect your AI agents to unstructured data in &lt;/span&gt;&lt;a href="https://cloud.google.com/storage"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Cloud Storage&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; via Model Context Protocol (MCP)? In &lt;/span&gt;&lt;a href="https://cloud.google.com/blog/topics/developers-practitioners/build-ai-agents-faster-with-gcs-google-cloud-storage-mcp-server"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;this blog&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;, learn about why would want to do that from three customer examples, then how to do it, choosing either a fully managed service, or a self-managed local server for more customization and control. &lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h4&gt;&lt;span style="vertical-align: baseline;"&gt;Research, reports and deep-dives&lt;/span&gt;&lt;/h4&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Report: &lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;According to an independent benchmark report, &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/kubernetes-engine/docs/concepts/about-gke-inference-gateway"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;GKE Inference Gateway&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; outperforms the next leading managed Kubernetes service with 15.7% higher throughput, 92.8% shorter wait times, and 62.6% lower inter-token latency. This performance can be attributed to its use of prefix caching, which optimizes LLM performance by storing the KV cache (activation states) of long, repetitive prompt prefixes. Learn more in the &lt;/span&gt;&lt;a href="https://cloud.google.com/blog/products/containers-kubernetes/gke-inference-gateway-prefix-caching-accelerates-ai-inference?e=0"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;blog&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;. &lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Architecture deep dive: &lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;A closer look at &lt;/span&gt;&lt;a href="https://discuss.google.dev/t/accelerate-tpu-model-loading-while-saving-ram-on-gke/374835" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;the cold start problem, this time for TPUs and GKE&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;, and how the Run:ai Model Streamer can help change the dynamic. &lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h4&gt;&lt;span style="vertical-align: baseline;"&gt;Customer and partner updates&lt;/span&gt;&lt;/h4&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Customer win:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Leveraging GKE, BigQuery, Cloud SQL, and Gemini Enterprise Agent Platform, &lt;/span&gt;&lt;a href="https://www.youtube.com/watch?v=x36QJ-QKRGg" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Pager Health is eliminating operational fragmentation to deliver a simplified, personalized U.S. healthcare experience&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; that transforms lives.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Customer win:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Trustpilot, the customer review platform, built a high-volume streaming pipeline using fine-tuned Gemma models with Dataflow and Gemini Enterprise Agent Platform running on cost-optimized A2 VMs using A100 GPUs, as well as optimized version of vLLM maintained by Gemini Enterprise Agent Platform.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;hr/&gt;
&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;May 2026&lt;/span&gt;&lt;/h3&gt;
&lt;h4&gt;&lt;span style="vertical-align: baseline;"&gt;Product, technology and tool updates&lt;/span&gt;&lt;/h4&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Product update:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/kubernetes-engine/docs/concepts/machine-learning/agent-sandbox"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;GKE Agent Sandbox&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; is now generally available.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;New open-source project:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; &lt;/span&gt;&lt;a href="https://github.com/agent-substrate/substrate" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Agent Substrate&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; is a new open-source project aimed at continuing to push the limits of agentic infrastructure density&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;New feature:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; &lt;/span&gt;&lt;a href="https://ai.google.dev/edge/ai-edge-portal" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Google AI Edge Portal&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;, a solution for testing and benchmarking on-device machine learning (ML) at scale, now supports benchmarking and debugging on-device LLMs. Read more &lt;/span&gt;&lt;a href="https://cloud.google.com/blog/products/ai-machine-learning/benchmark-llms-on-device-with-ai-edge-portal?e=48754805"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;here&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;. &lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Product deep dive: &lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;We went &lt;/span&gt;&lt;a href="https://cloud.google.com/blog/products/storage-data-transfer/cloud-storage-rapid-turbocharges-object-storage-for-ai-analytics?e=48754805"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;into depth about Cloud Storage Rapid&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;, a new family of high-performance storage offerings for AI workloads. At launch, offerings include Rapid Bucket (formerly Rapid Storage), a high-performance zonal object storage offering, and Rapid Cache (formerly Anywhere Cache), which accelerates reads on-demand and colocates compute and data for workloads in existing buckets. &lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h4&gt;&lt;span style="vertical-align: baseline;"&gt;Research, reports and deep dives&lt;/span&gt;&lt;/h4&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Architecture deep dive: &lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;Google Global Infrastructure VP Bikash Koley and Engineering Fellow Arjun Singh provide a high-level overview of &lt;/span&gt;&lt;a href="https://cloud.google.com/blog/products/networking/data-center-and-global-networks-built-for-ai-era"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;the challenges that AI workloads pose to network infrastructure&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;, and discuss the deep enhancements we’ve made to our data center fabrics, WAN, and global networks to better support them. &lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Architecture deep dive: &lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;We unveiled a &lt;/span&gt;&lt;a href="https://cloud.google.com/blog/products/compute/cluster-reliability-for-trillion-parameter-models-on-tpus?e=48754805"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;new cluster-level reliability model&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; for developing frontier AI models on TPUs, ditching instance-level reliability &lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h4&gt;&lt;span style="vertical-align: baseline;"&gt;Customer and partner updates&lt;/span&gt;&lt;/h4&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Customer win:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Visual media provider &lt;/span&gt;&lt;a href="https://cloud.google.com/blog/products/infrastructure/how-imgix-processes-8-billion-images-daily-with-g4-vms-powered-by-nvidia-blackwell?e=48754805"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Imgix serves more than 8 billion images and videos from AI Hypercomputer&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; equipped with G4 VMs powered by NVIDIA RTX PRO 6000 Blackwell GPUs.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;&lt;/div&gt;</description><pubDate>Mon, 31 Aug 2026 16:00:00 +0000</pubDate><guid>https://cloud.google.com/blog/topics/ai-infrastructure/whats-new-in-ai-infrastructure-this-month/</guid><category>AI &amp; Machine Learning</category><category>Containers &amp; Kubernetes</category><category>Compute</category><category>Networking</category><category>Storage &amp; Data Transfer</category><category>AI infrastructure</category><media:content height="540" url="https://storage.googleapis.com/gweb-cloudblog-publish/images/Whats_new_in_AI_infrastructure.max-600x600.jpg" width="540"></media:content><og xmlns:og="http://ogp.me/ns#"><type>article</type><title>What’s new in AI infrastructure and orchestration in August</title><description></description><image>https://storage.googleapis.com/gweb-cloudblog-publish/images/Whats_new_in_AI_infrastructure.max-600x600.jpg</image><site_name>Google</site_name><url>https://cloud.google.com/blog/topics/ai-infrastructure/whats-new-in-ai-infrastructure-this-month/</url></og><author xmlns:author="http://www.w3.org/2005/Atom"><name>Alex Barrett</name><title>Editor, Google Cloud blog</title><department></department><company></company></author></item><item><title>Reimagining work: How Pythian’s internal AI playbook delivers customer ROI</title><link>https://cloud.google.com/blog/topics/startups/how-pythians-internal-ai-playbook-delivers-customer-roi/</link><description>&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;When &lt;/span&gt;&lt;a href="https://www.pythian.com/" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Pythian&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; rolled out Google Cloud’s &lt;/span&gt;&lt;a href="https://cloud.google.com/gemini-enterprise"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Gemini Enterprise&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; across our 500-person company in 27 countries, the goal was simple: use our own company as a proving ground to discover how enterprise AI actually delivers ROI.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;What we found changed our strategy entirely.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Since the rollout of Gemini Enterprise and our previous enterprise AI deployments, Pythian observed firsthand why so many enterprise AI initiatives stall out or fail. &lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Most organizations trap themselves in a tool-centric mindset — buying licenses, making tools broadly available, and assuming value will naturally follow. They get stuck chasing "nickel and dime" micro-efficiencies (like saving 5 minutes per user) while missing structural, high-ROI workflow transformations. Compounding the problem, even when custom agents are built, they frequently stall in pilot mode or break down in production because teams lack the operational capability to manage AI model drift, agent lifecycles, and ongoing observability.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;To solve this, we engineered the Pythian AI Operating Model — a multifaceted, end-to-end framework designed to take enterprise AI from high-level strategy all the way into sustained production. While our dual center of excellence (COE) serves as the core execution muscle, it is the application of the entire framework, from Field CTO strategy and tooling deployment to the dual COE and XOps, that consistently unlocks million-dollar outcomes.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;By proving this complete model internally first, Pythian drove a&lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt; &lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;3x&lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt; &lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;surge in active user engagement and cut our database incident resolution times by 80%.&lt;/span&gt;&lt;/p&gt;
&lt;h2&gt;&lt;strong style="vertical-align: baseline;"&gt;The four pillars of the Pythian AI operating model&lt;/strong&gt;&lt;/h2&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;To move past the common failure points of enterprise AI, our framework consolidates strategy, execution, and operations into a single continuous loop:&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;Field CTO strategy  ──&amp;gt;  tooling deployment  ──&amp;gt;  dual COE execution  ──&amp;gt;  production XOps&lt;/strong&gt;&lt;/p&gt;
&lt;ol&gt;
&lt;li aria-level="1" style="list-style-type: decimal; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Field CTO strategy and governance:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Generative AI is arguably the most academically challenging architectural shift in IT history. Led by former C-suite tech leaders, our Field CTO practice provides executive advisory to establish steering committees and clear value metrics. The team audits operations using 16 horizontal agentic patterns (like automated document processing and runbook creation) to build a prioritized backlog of high-ROI use cases &lt;/span&gt;&lt;span style="font-style: italic; vertical-align: baseline;"&gt;before&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt; development starts.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: decimal; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Tooling and platform deployment:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; The team establishes a secure, production-grade foundation on platforms like Gemini Enterprise and connects AI directly into CRMs, ERPs, and database estates to ground models in real corporate context.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: decimal; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;The dualCOE:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; This execution muscle is split into two specialized engines:&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;ul&gt;
&lt;li aria-level="2" style="list-style-type: circle; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;People productivity COE:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; This group handles adoption and change management. Instead of expecting non-technical teams (like HR or Procurement) to build its own agents, this COE builds no-code agents &lt;/span&gt;&lt;span style="font-style: italic; vertical-align: baseline;"&gt;for&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt; them, focusing entirely on enablement.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="2" style="list-style-type: circle; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Process productivity COE:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; This team engineers deep, custom-coded AI agents and complex agentic workflows that integrate into core data platforms for autonomous operations.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;li aria-level="1" style="list-style-type: decimal; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;XOps (AI production management):&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; While deploying an agent is 20% of the journey,  &lt;/span&gt;&lt;span style="font-style: italic; vertical-align: baseline;"&gt;maintaining&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt; accuracy in production is 80%. Because AI models and prompt structures naturally drift over time, this XOps practice provides the continuous monitoring, prompt tuning, and model observability needed to keep agents performing without breaking core workflows.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;The difference between chasing minor, scattered efficiencies and driving structural enterprise ROI comes down to how you align your operating strategy:&lt;br/&gt;&lt;br/&gt;&lt;/span&gt;&lt;/p&gt;
&lt;div align="left"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;&lt;table&gt;&lt;colgroup&gt;&lt;col/&gt;&lt;col/&gt;&lt;col/&gt;&lt;/colgroup&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Alignment element&lt;/span&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Tool-centric approach&lt;/span&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Pythian AI operating model&lt;/span&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Primary metric&lt;/span&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Individual minutes saved per user&lt;/span&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;High-impact workflow reimagination and ROI&lt;/span&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Operational focus&lt;/span&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Broad, unguided tool availability&lt;/span&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Prioritized backlog via 16 agentic patterns&lt;/span&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Execution muscle&lt;/span&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Ad-hoc user experimentation&lt;/span&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Dual COE (people and process productivity)&lt;/span&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Production lifecycle&lt;/span&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Unmonitored static deployments&lt;/span&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Active XOps (Continuous accuracy and drift management)&lt;/span&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;h2&gt;&lt;strong style="vertical-align: baseline;"&gt;Real-world impact: from database ops to global supply chains&lt;/strong&gt;&lt;/h2&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Whether managing 70 manufacturing plants or 30,000 enterprise databases, AI succeeds when tied to structural, high-value workflows:&lt;/span&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Pythian “as a customer:”&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Across 15,000 monthly database tickets, our Process COE deployed an agentic workflow that reads tickets, searches knowledge bases, and auto-generates mini runbooks before an engineer touches them. The result was slashed mean time to resolution by 80% and tripled active user engagement&lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;.&lt;/strong&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Knowledge management customer:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; We deployed autonomous IT support agents across 10,000 consultants. As a result, we were able to automate 10% of 20,000 annual IT tickets into "no-touch" resolutions, saving 1,000,000+ operational hours&lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;.&lt;/strong&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Supply chain customer:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; By building custom agentic supply chain tools on Gemini Enterprise, we compressed forecast-matching cycles from weeks down to 2–3 days across 70 global manufacturing sites&lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;.&lt;/strong&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Retail customer:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; We combined &lt;/span&gt;&lt;a href="https://cloud.google.com/gemini-enterprise/agents"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Gemini Agentic AI&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; and computer vision to automate store product onboarding. As a result, we transformed a 20-minute manual task into a multi-second flow&lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;.&lt;/strong&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h2&gt;&lt;strong style="vertical-align: baseline;"&gt;Ready to build your AI operating model?&lt;/strong&gt;&lt;/h2&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Scaling AI demands more than tool-level experimentation. It also requires an end-to-end AI operating model. Learn how Pythian pairs with Google Cloud to operationalize strategy, streamline XOps, and fast-track your Gemini Enterprise journey.&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;</description><pubDate>Thu, 27 Aug 2026 16:00:00 +0000</pubDate><guid>https://cloud.google.com/blog/topics/startups/how-pythians-internal-ai-playbook-delivers-customer-roi/</guid><category>AI &amp; Machine Learning</category><category>Customers</category><category>Startups</category><media:content height="540" url="https://storage.googleapis.com/gweb-cloudblog-publish/images/pythian-ai-framework-blog-header.max-600x600.png" width="540"></media:content><og xmlns:og="http://ogp.me/ns#"><type>article</type><title>Reimagining work: How Pythian’s internal AI playbook delivers customer ROI</title><description></description><image>https://storage.googleapis.com/gweb-cloudblog-publish/images/pythian-ai-framework-blog-header.max-600x600.png</image><site_name>Google</site_name><url>https://cloud.google.com/blog/topics/startups/how-pythians-internal-ai-playbook-delivers-customer-roi/</url></og><author xmlns:author="http://www.w3.org/2005/Atom"><name>Paul Lewis</name><title>Chief Technology Officer, Pythian</title><department></department><company></company></author><author xmlns:author="http://www.w3.org/2005/Atom"><name>Vanessa Simmons</name><title>SVP, Business Development, Pythian</title><department></department><company></company></author></item><item><title>FinOps for the AI era: New flexible billing and cost controls for agents</title><link>https://cloud.google.com/blog/products/ai-machine-learning/flexible-billing-and-cost-controls-for-agents-on-google-cloud/</link><description>&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;strong&gt;&lt;span style="vertical-align: baseline;"&gt;Editor's note:&lt;/span&gt;&lt;/strong&gt;&lt;em&gt;&lt;span style="vertical-align: baseline;"&gt; A product image was updated after initial publication.&lt;/span&gt;&lt;/em&gt;&lt;/p&gt;
&lt;hr/&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;As AI takes on more complex work, business leaders face a new challenge: enabling rapid innovation using agents while protecting their margins and budgets. To get a real return on AI, financial operations (FinOps) and cost management must evolve alongside technology, giving you clear visibility, proactive cost controls, and flexible payment models that fit your needs. &lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;That’s why today we’re introducing &lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;expanded billing flexibility and new cost management tools for agent workloads &lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;across Gemini Enterprise and developer tools like &lt;/span&gt;&lt;a href="https://cloud.google.com/blog/products/ai-machine-learning/expanding-google-antigravity-for-enterprise-customers"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Google Antigravity in Gemini Enterprise &lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;and &lt;/span&gt;&lt;a href="http://d.android.com/gemini-in-android" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Android Studio&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;.&lt;/span&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Flexible payment options:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; You can mix our existing, predictable per-user seat subscriptions with a &lt;/span&gt;&lt;a href="https://cloud.google.com/gemini-enterprise#gemini-enterprise-app-editions"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;new pay-as-you-go option&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; in Gemini Enterprise app that lets you run agent workloads without hitting quota limits mid-task.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Developer access, one place to manage your AI: &lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;Google Antigravity and Android Studio AI use is now included in your Gemini Enterprise subscription (available for select customers and rolling out broadly soon), giving your developers more without giving you more to manage. Usage across Antigravity, the platform, and the app rolls up into a single view instead of separate licenses and billing silos.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Pay less as your usage grows:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; If your AI workloads are steady or climbing, &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/docs/cuds-flexible-savings-plans"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Flexible Savings Plans&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; let you commit to a monthly spend you're comfortable with and take 10–20% off your token costs — no minimums, no maximums, and no new billing silo to manage.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Consolidated spend guardrails: &lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;You can now set hard monthly caps on AI spend and projects, estimate agent runtime costs, and catch sudden budget spikes before they hit your invoice.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h2&gt;&lt;span style="vertical-align: baseline;"&gt;Give your teams flexibility without losing control over spend in Gemini Enterprise&lt;/span&gt;&lt;/h2&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Every organization operates differently. Even within the same business, no two teams consume AI &lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;in the same way&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;. Your business users might rely on steady, everyday productivity tools. Meanwhile, your technical teams might run AI agent workloads in bursts. &lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;To help align costs with how work actually gets done, you can combine these payment and licensing choices and features across Gemini Enterprise:&lt;br/&gt;&lt;br/&gt;&lt;/span&gt;&lt;/p&gt;
&lt;div align="left"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;&lt;table&gt;&lt;colgroup&gt;&lt;col/&gt;&lt;col/&gt;&lt;col/&gt;&lt;/colgroup&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;Option&lt;/strong&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;How it works&lt;/strong&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;Why it helps optimize spend&lt;/strong&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;Gemini Enterprise app per-user seat subscription&lt;/strong&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;You pay a fixed monthly fee per user, which includes daily quota pools that are shared across your entire project.&lt;/span&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;Predictable budgeting.&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; It provides finance teams with a clear, steady monthly baseline for teams with consistent daily productivity needs.&lt;/span&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;[New] Gemini Enterprise app pay-as-you-go consumption edition&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="font-style: italic; vertical-align: baseline;"&gt;*available for select customers and rolling out broadly soon&lt;/span&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;There is no upfront commitment or base subscription fee, meaning you pay strictly for the compute and tokens your teams consume at standard model API rates.&lt;/span&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;Only pay for what you use.&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Your spend scales up and down automatically with real usage, ensuring you never pay for empty seats when project demand dips.&lt;/span&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;[New for Antigravity in Gemini Enterprise] Consolidated pooled quotas&lt;/strong&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Daily usage allowances are pooled project-wide, letting business apps, developer tools, and custom agents draw from the same shared quota. Pooled quota is always exhausted first, and admins can control if overages are allowed, at which point it’s charged at pay-as-you-go rates. &lt;/span&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;Maximized resource usage:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Unused daily allowances from business users automatically absorb heavy developer or custom API agent demands, so no quota allowance goes to waste.&lt;/span&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;[Coming soon] &lt;/strong&gt;&lt;strong style="vertical-align: baseline;"&gt;Deferred execution pricing&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="font-style: italic; vertical-align: baseline;"&gt;*available for select workloads soon&lt;/span&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Mark eligible agent workloads as deferred, and our intelligent scheduler in the Gemini Enterprise Agent Platform runs them during off-peak capacity windows.&lt;/span&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;Substantial discounts for work that can wait: &lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;AI workloads can run on separate, off-peak capacity, you pay up to half the inference cost and bypass standard quota limits entirely – letting you run substantially more agentic volume under the same budget.&lt;/span&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;h3&gt;&lt;strong style="vertical-align: baseline;"&gt;Equip developers with advanced agentic tooling under a single Gemini Enterprise subscription&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;We’re rolling out access to &lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;Google Antigravity in Gemini Enterprise&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;, an agent-first developer platform that brings powerful agentic coding and agent-building capabilities to technical teams, included with Gemini Enterprise subscriptions for &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/gemini/enterprise/docs/ai-developer-tools-overview"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;eligible customers&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;. In addition, Android developers can leverage the Google Antigravity quota included in their Gemini Enterprise subscriptions natively in &lt;/span&gt;&lt;a href="http://d.android.com" rel="noopener" target="_blank"&gt;&lt;strong style="text-decoration: underline; vertical-align: baseline;"&gt;Android Studio&lt;/strong&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;, the agentic IDE for professional Android development.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;To be more efficient with agentic coding costs, we are pooling developer tools quota included in each Gemini Enterprise subscription and making it available across the whole Google Cloud project so your teams can benefit from the capacity you’re already purchasing. Your developers get access to advanced agentic tools, while you maintain centralized governance and control.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;For a closer look into what’s new with Antigravity in Gemini Enterprise and how customers are putting it to work in production, take a look at our &lt;/span&gt;&lt;a href="https://cloud.google.com/blog/products/ai-machine-learning/expanding-google-antigravity-for-enterprise-customers"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;deep-dive&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;.&lt;/span&gt;&lt;/p&gt;
&lt;h3&gt;&lt;strong style="vertical-align: baseline;"&gt;Budget smarter with Gemini Enterprise Flexible Savings Plans (FSPs) &lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;If your organization has steady or growing AI workloads, Gemini Enterprise Flexible Savings Plans offer a simple, spend-based commitment model across Gemini Enterprise usage. FSPs are designed to lower token costs while keeping budgets flexible:&lt;/span&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Programmatic savings: &lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;Receive 10% off for 1-year or 20% off for 3-year commitments for monthly spending across Gemini Enterprise.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Tailored to your pace&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;: With no minimum or maximum spend requirements, you can determine a monthly commitment that fits your current traffic and make adjustments as your usage increases over time. &lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Enterprise Agreement (EA) friendly:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; FSP spend seamlessly draws down against your existing Google Cloud EA, giving lines of business dedicated budget control without fragmenting your broader cloud commitments.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;a href="https://cloud.google.com/gemini-enterprise-agent-platform/generative-ai/pricing"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Gemini Enterprise Flexible Savings Plans&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; are already available for self-serve customers and customers on enterprise agreements.&lt;/span&gt;&lt;/p&gt;
&lt;h3&gt;&lt;strong style="vertical-align: baseline;"&gt;Give your teams the freedom to build while maintaining financial discipline&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;As a leader, your goal isn't to restrict the potential value of AI  – it's to remove the financial and operational risk that you face without managed AI costs. You should be able to give engineering, marketing, and operational teams the freedom to innovate with agents, but you should also have the visibility to trust what those agents are doing and the safety nets to protect your budget.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;To bridge this gap, we've built robust, native governance tooling directly into the Google Cloud Billing Console around three simple goals:&lt;/span&gt;&lt;/p&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;1. Plan before you scale: &lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;The &lt;/span&gt;&lt;a href="https://cloud.google.com/products/calculator"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Google Cloud Pricing Calculator&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; lets you estimate anticipated costs in Gemini Enterprise across per-user licenses, developer tools, and background agent runtimes. It gives you the numbers you need to build clear business cases upfront before project work begins&lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;.&lt;/strong&gt;&lt;/p&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;2. Enforce boundaries without micromanaging spend: &lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;Instead of spending time tracking daily usage variations across project teams, let these tools do the monitoring for you:&lt;/span&gt;&lt;/p&gt;
&lt;ol start="2"&gt;
&lt;ul&gt;
&lt;li aria-level="2" style="list-style-type: circle; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Early anomaly detection: &lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;If a project’s AI spending trends higher than normal, the system flags the deviation with root cause analysis and pinpoints the top 3 SKUs driving the increase so you can see exactly what changed.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;/ol&gt;&lt;/div&gt;
&lt;div class="block-image_full_width"&gt;






  
    &lt;div class="article-module h-c-page"&gt;
      &lt;div class="h-c-grid"&gt;
  

    &lt;figure class="article-image--large
      
      
        h-c-grid__col
        h-c-grid__col--6 h-c-grid__col--offset-3
        
        
      "
      &gt;

      
      
        
        &lt;img
            src="https://storage.googleapis.com/gweb-cloudblog-publish/images/1_Jul22_Anomalies_Image1.max-1000x1000.png"
        
          alt="1 Jul22_Anomalies_Image1"&gt;
        
        &lt;/a&gt;
      
        &lt;figcaption class="article-image__caption "&gt;&lt;p data-block-key="eqcz8"&gt;Billing Console showing an Early Anomaly alert with the Root Cause Analysis (RCA) breakdown highlighting the driving SKUs&lt;/p&gt;&lt;/figcaption&gt;
      
    &lt;/figure&gt;

  
      &lt;/div&gt;
    &lt;/div&gt;
  




&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;ul&gt;
&lt;li style="list-style-type: none;"&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong style="vertical-align: baseline;"&gt;Project-level spend caps:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; When a project needs defined financial boundaries, you can set a firm monthly spend limit directly in the Google Cloud Billing Console. If a project hits its limit, the agent's API calls temporarily pause – protecting your budget without affecting the rest of your production infrastructure. Automated email alerts at 50%, 80% and 100% of the budget keep you informed of your progress against the spend limit. &lt;/span&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;ul&gt;
&lt;li style="list-style-type: none;"&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong style="vertical-align: baseline;"&gt;Overage controls:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; If a spend cap triggers, you can choose to resume work with a single click in the console. Alternatively, if your priority is continuous operation, you can turn on overages so excess usage smoothly transitions to consumption rates, which can draw directly against your FSP to keep overage unit costs heavily discounted.&lt;/span&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;&lt;/div&gt;
&lt;div class="block-image_full_width"&gt;






  
    &lt;div class="article-module h-c-page"&gt;
      &lt;div class="h-c-grid"&gt;
  

    &lt;figure class="article-image--large
      
      
        h-c-grid__col
        h-c-grid__col--6 h-c-grid__col--offset-3
        
        
      "
      &gt;

      
      
        
        &lt;img
            src="https://storage.googleapis.com/gweb-cloudblog-publish/images/3_PAYG_Overage_Enabled.max-1000x1000.png"
        
          alt="3 PAYG Overage Enabled"&gt;
        
        &lt;/a&gt;
      
        &lt;figcaption class="article-image__caption "&gt;&lt;p data-block-key="eqcz8"&gt;Enabling overage pay-as-you-go for a project.&lt;/p&gt;&lt;/figcaption&gt;
      
    &lt;/figure&gt;

  
      &lt;/div&gt;
    &lt;/div&gt;
  




&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;3. Get visibility into business value:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Use centralized billing reports paired with the FinOps agent to generate natural-language cost insight summaries of where your budget went, making it simple to show ROI to leadership.&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-image_full_width"&gt;






  
    &lt;div class="article-module h-c-page"&gt;
      &lt;div class="h-c-grid"&gt;
  

    &lt;figure class="article-image--large
      
      
        h-c-grid__col
        h-c-grid__col--6 h-c-grid__col--offset-3
        
        
      "
      &gt;

      
      
        
        &lt;img
            src="https://storage.googleapis.com/gweb-cloudblog-publish/images/Billing_overage_-_dashboard_-_new_afternoo.max-1000x1000.png"
        
          alt="Cost overview FinOps"&gt;
        
        &lt;/a&gt;
      
        &lt;figcaption class="article-image__caption "&gt;&lt;p data-block-key="a5o2c"&gt;AI spending reporting in Google Cloud Console&lt;/p&gt;&lt;/figcaption&gt;
      
    &lt;/figure&gt;

  
      &lt;/div&gt;
    &lt;/div&gt;
  




&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;Go deeper with AI cost optimization&lt;/span&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;To build a full-stack FinOps strategy that optimizes the cost, latency, and performance of your models and infrastructure, explore our detailed architecture specifications and frameworks:&lt;/span&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;span style="vertical-align: baseline;"&gt;&lt;a href="https://cloud.google.com/blog/topics/ai-infrastructure/best-practices-for-dynamic-capacity-management"&gt;&lt;strong style="text-decoration: underline; vertical-align: baseline;"&gt;How to outsmart infrastructure constraints with dynamic capacity management&lt;/strong&gt;&lt;/a&gt;&lt;strong&gt;&lt;span style="vertical-align: baseline;"&gt;:&lt;/span&gt;&lt;/strong&gt; Discover how to optimize your compute investments with capabilities in Google Kubernetes Engine and Google Compute Engine that automatically schedule and reallocate resources to avoid interruptions, over-provisioning, and over-reliance on any one hardware configuration.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;a href="https://cloud.google.com/blog/products/ai-machine-learning/expanding-google-antigravity-for-enterprise-customers"&gt;&lt;strong style="text-decoration: underline; vertical-align: baseline;"&gt;Expanding Google Antigravity for Enterprise Customers&lt;/strong&gt;&lt;/a&gt;&lt;strong style="vertical-align: baseline;"&gt;:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Read our developer tooling deep-dive to see how technical teams are accelerating software delivery with agent-first workflows.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;&lt;a href="https://cloud.google.com/transform/gemini-enterprise-optimize-ai-token-spend"&gt;&lt;strong style="text-decoration: underline; vertical-align: baseline;"&gt;What sports cars can teach us about optimizing AI spend&lt;/strong&gt;&lt;/a&gt;&lt;strong style="vertical-align: baseline;"&gt;: &lt;/strong&gt;&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;More tokens doesn't always mean better AI. Read our conversation with Mike Clark, Director of Product Management for Gemini Enterprise Agent Platform, on how to balance horsepower with efficiency and get the highest return out of every dollar you spend on AI. &lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;a href="https://docs.cloud.google.com/gemini-enterprise-agent-platform/models/provisioned-throughput"&gt;&lt;strong style="text-decoration: underline; vertical-align: baseline;"&gt;Protection during usage spikes&lt;/strong&gt;&lt;/a&gt;&lt;strong style="vertical-align: baseline;"&gt;:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Your heavy workloads can surge during peak hours without forcing you to pay for expensive, dedicated infrastructure that sits idle the rest of the time. As your AI usage grows, Gemini models can automatically scale on demand without hitting artificial rate limits – processing up to 50 million tokens per minute. Read more about Provisioned Throughput.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;&lt;/div&gt;</description><pubDate>Wed, 26 Aug 2026 13:30:00 +0000</pubDate><guid>https://cloud.google.com/blog/products/ai-machine-learning/flexible-billing-and-cost-controls-for-agents-on-google-cloud/</guid><category>Cost Management</category><category>AI &amp; Machine Learning</category><media:content height="540" url="https://storage.googleapis.com/gweb-cloudblog-publish/images/FinOps_for_the_AI_era_.max-600x600.jpg" width="540"></media:content><og xmlns:og="http://ogp.me/ns#"><type>article</type><title>FinOps for the AI era: New flexible billing and cost controls for agents</title><description></description><image>https://storage.googleapis.com/gweb-cloudblog-publish/images/FinOps_for_the_AI_era_.max-600x600.jpg</image><site_name>Google</site_name><url>https://cloud.google.com/blog/products/ai-machine-learning/flexible-billing-and-cost-controls-for-agents-on-google-cloud/</url></og><author xmlns:author="http://www.w3.org/2005/Atom"><name>Michael Gerstenhaber</name><title>VP, Product Management, Gemini Enterprise</title><department></department><company></company></author><author xmlns:author="http://www.w3.org/2005/Atom"><name>Pravir Gupta</name><title>VP, Google Business Platform</title><department></department><company></company></author></item><item><title>Now introducing Gemini Enterprise for Legal</title><link>https://cloud.google.com/blog/products/ai-machine-learning/introducing-gemini-enterprise-for-legal/</link><description>&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Few professions are as exacting as the practice of law. A team reviewing a contract or building a case works inside strictly privileged information, firm-specific playbooks, and a body of law that changes constantly. The work thrives on nuanced, professional judgment — and the systems supporting it inherit real obligations: ethical walls that cannot be crossed, matter permissions that cannot be flattened, and a duty of confidentiality that does not bend for convenience.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;General-purpose AI, however capable, does not meet that standard on its own. Foundational model intelligence is necessary. For legal work, it is nowhere near sufficient.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;What makes the difference is the system built around the model: &lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;skills that enhance a firm's own expertise, connections into the systems where matters actually live, agents that complete work rather than return suggestions, and an open ecosystem to extend all of it — with governance running underneath all four&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;. Each is valuable alone. Only in combination do they produce something a firm or a legal department can put into production and actually rely on.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Today we're bringing that to legal practice with Gemini Enterprise for Legal, part of our new suite of purpose-built industry solutions.&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-video"&gt;



&lt;div class="article-module article-video "&gt;
  &lt;figure&gt;
    &lt;a class="h-c-video h-c-video--marquee"
      href="https://youtube.com/watch?v=ct5JDCb0-BA"
      data-glue-modal-trigger="uni-modal-ct5JDCb0-BA-"
      data-glue-modal-disabled-on-mobile="true"&gt;

      
        

        &lt;div class="article-video__aspect-image"
          style="background-image: url(https://storage.googleapis.com/gweb-cloudblog-publish/images/image2_SBQjmnB.max-1000x1000.png);"&gt;
          &lt;span class="h-u-visually-hidden"&gt;Gemini Enterprise for Legal&lt;/span&gt;
        &lt;/div&gt;
      
      &lt;svg role="img" class="h-c-video__play h-c-icon h-c-icon--color-white"&gt;
        &lt;use xlink:href="#mi-youtube-icon"&gt;&lt;/use&gt;
      &lt;/svg&gt;
    &lt;/a&gt;

    
  &lt;/figure&gt;
&lt;/div&gt;

&lt;div class="h-c-modal--video"
     data-glue-modal="uni-modal-ct5JDCb0-BA-"
     data-glue-modal-close-label="Close Dialog"&gt;
   &lt;a class="glue-yt-video"
      data-glue-yt-video-autoplay="true"
      data-glue-yt-video-height="99%"
      data-glue-yt-video-vid="ct5JDCb0-BA"
      data-glue-yt-video-width="100%"
      href="https://youtube.com/watch?v=ct5JDCb0-BA"
      ng-cloak&gt;
   &lt;/a&gt;
&lt;/div&gt;

&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;p style="text-align: center;"&gt;&lt;em&gt;Bringing Gemini Enterprise to your legal practice&lt;/em&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;Four components of Gemini Enterprise for Legal&lt;/span&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Developed alongside industry leaders, Gemini Enterprise for Legal provides an integrated, fully governed environment configured for rapid deployment across firms and corporate legal departments:&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;1. Purpose-built skills for legal work.&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Skills are reusable packages of instructions and context, designed by domain experts, that teach an agent to run a specialized task while enforcing your firm's playbooks, citation rules, and house style. They cover contract review and redlining, playbook creation, regulatory horizon scanning, legal research, DSAR fulfillment, and more — and they are where a firm's institutional knowledge becomes something the platform can execute rather than something a partner has to re-explain.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;2. Connections to trusted systems and data.&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Secure MCP connectors link agents to the document management systems, case repositories, research services, and industry applications legal teams already rely on — inheriting each platform's existing user permissions and access controls rather than working around them.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;3. Agents that act within the data.&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Skills and connections come together in agents that carry work through: pre-built agents from Google and leading legal software providers deploy out of the box. Specialized agents handle legal and policy research, regulatory screening, and contract drafting — bringing deep legal expertise onto a platform with centralized governance.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;4. An open partner ecosystem.&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Every firm and legal department practices differently. Partnerships with global systems integrators and legal-tech specialists — &lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;Accenture&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;, &lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;Deloitte&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;, &lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;Devoteam&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;, &lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;Factor Law&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;, &lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;KPMG&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;, &lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;Tribe.ai&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;, &lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;Valtech&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;, &lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;Zazmic&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;, &lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;Zencore&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;, and &lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;66degrees&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; — let organizations customize, integrate, and scale across complex enterprise architectures without vendor lock-in.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;Running underneath: a governed control plane.&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; A single dashboard for legal IT and risk teams that natively &lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;enforces&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; security policies (VPC, CMEK), maintains private data isolation, and holds every output to verifiable grounding with traceable citations.&lt;/span&gt;&lt;/p&gt;
&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;Unlocking high-value workflows with domain-specific skills&lt;/span&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Gemini Enterprise for Legal shifts AI from passive querying to agentic execution, automating high-volume, precision-critical workflows such as:&lt;/span&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Proactive regulatory horizon scanning:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Keeps legal and compliance teams ahead of global mandates by autonomously tracking legislative updates, court dockets, and supervisory bodies. It cross-references emerging changes against enterprise policies to flag exposure gaps and generate updated policy drafts for immediate practitioner review.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Automating data discovery and DSAR response:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Modernizes privacy workflows by compiling personal data across fragmented enterprise systems in seconds. It eliminates the manual toil of Data Subject Access Requests (DSARs), and allows for adherence to regulatory timelines while minimizing operational risk.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Accelerating contract review and negotiation:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Compresses turnaround times for inbound vendor agreements, NDAs, and complex M&amp;amp;A documentation by benchmarking terms against enterprise playbooks. It surfaces high-risk clauses and potential exposure, enabling attorneys to focus on strategic negotiation and high-value judgment.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Building and updating contracting playbooks:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Transforms legacy agreement archives into dynamic, actionable playbooks instantly. It automatically extracts key terms, fallback positions, and institutional knowledge to maintain portfolio-wide term consistency and lower negotiation variance across the enterprise.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Redacting documents for motions to seal:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Eliminates the manual burden of preparing court filings and redacting legal documents. It intelligently identifies sensitive terms and PII for rapid practitioner confirmation, dramatically accelerating filing timelines while safeguarding confidentiality.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Drafting NDA documents:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Elevates contract creation through structural fidelity validation that enforces firm standards and logical document hierarchies. It allows legal teams to rapidly generate and evolve non-disclosure agreements with complete formatting confidence and minimal review overhead.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;Open ecosystem of connectors across the legal technology stack&lt;/span&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Legal work is only as good as its sources, and legal data carries permissions that have to travel with it. Gemini Enterprise for Legal connects directly to core legal systems via secure &lt;a href="https://cloud.google.com/gemini-enterprise/connectors?e=48754805"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;MCP connectors&lt;/span&gt;&lt;/a&gt;. Crucially, access is bound by existing role-based access controls, document-level permissions, and trusted data controls inherited from document management and ediscovery systems.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;&lt;span style="vertical-align: baseline;"&gt;Productivity and collaboration:&lt;/span&gt;&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Google Workspace:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Connects seamlessly with Google Docs, Gmail, Drive, and Sheets to analyze matter communications, correspondence, and surface internal files while enforcing enterprise access controls.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Microsoft 365:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Integrates directly with Word, Outlook, and SharePoint to triage inquiries, redlines, and securely ground work product across emails and matter folders without breaking workflow context. &lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;&lt;span style="vertical-align: baseline;"&gt;Document management:&lt;/span&gt;&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;iManage:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; &lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;Gives Gemini Enterprise for Legal permission-bound, auditable access to governed iManage content, including matter history, documents, and institutional knowledge, eliminating the need for bulk exports or custom integrations.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;NetDocuments&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;: &lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;Enables Gemini Enterprise to search and analyze an organization's knowledge and expertise while preserving each user's existing permissions and ethical walls. Source documents never leave the governed NetDocuments environment. &lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;&lt;span style="vertical-align: baseline;"&gt;Contract lifecycle and execution:&lt;/span&gt;&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Docusign:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Integrates agreement metadata, active approval workflows, and contract repositories to surface obligations, track renewal dates, and streamline drafting-to-execution lifecycles.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;&lt;span style="vertical-align: baseline;"&gt;E-discovery and litigation intelligence:&lt;/span&gt;&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Everlaw&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;: &lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;Connects Gemini Enterprise to litigation and investigations evidence in Everlaw, allowing legal teams to search and analyze their data, uncover case insights, and build timelines directly in Gemini Enterprise, with responses grounded in the underlying documents and access governed by each user’s existing Everlaw permissions.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;RelativityOne:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Allows legal teams to stand up workspaces, organize case data, and manage operations within a secure perimeter. &lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;&lt;span style="vertical-align: baseline;"&gt;Primary law, research, and public dockets:&lt;/span&gt;&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Thomson Reuters HighQ:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; &lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;Connects Gemini Enterprise for Legal with HighQ, helping legal teams securely access and reference relevant HighQ content within their workflows. &lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Free Law Project’s CourtListener.com&lt;/strong&gt;&lt;strong style="vertical-align: baseline;"&gt;:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Provides access to millions of federal and state court opinions, PACER dockets, judicial profiles, and oral arguments.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Courtroom5:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; &lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;Delivers jurisdiction-aware civil litigation datasets, procedural rules, and deadline calculation logic.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;&lt;span style="vertical-align: baseline;"&gt;Specialized legal AI and intellectual property:&lt;/span&gt;&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Harvey:&lt;/strong&gt;&lt;span style="font-style: italic; vertical-align: baseline;"&gt; &lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;Bridges Harvey’s legal reasoning intelligence into Gemini Enterprise, supporting complex legal reasoning and research across Vault projects. &lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Solve Intelligence:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Links Gemini Enterprise to worldwide patent and non-patent literature, SEP technical standards, and prior art databases for patent drafting and claim charting.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Legora: &lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;Agentic operating system for legal work, supporting lawyers in research, review, and drafting across complex matters &lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;Third-party agents and implementation partners &lt;/span&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Every firm and legal department practices differently. Through our open platform, organizations can deploy pre-built partner agents or collaborate with systems integrators to scale custom capabilities:&lt;/span&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Deloitte: &lt;/strong&gt;&lt;a href="https://console.cloud.google.com/marketplace/product/us-con-gcp-sbx-0000427-020625/deloitte-contractsummarize-pro-agent" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Contract Summarize Pro Agent&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; that synthesizes complex contracts into clear summaries for rapid insight and informed decision-making. &lt;/span&gt;&lt;a href="https://console.cloud.google.com/marketplace/product/us-con-gcp-sbx-0000427-020625/deloitte-clause-guard-contract-redlining-agent" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Clause Guard&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; contract redlining agent to accelerate turnaround times, and minimize risk in contract management. &lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;a href="https://console.cloud.google.com/marketplace/product/eudia/eudia-knowledgebase"&gt;&lt;strong style="text-decoration: underline; vertical-align: baseline;"&gt;Eudia Knowledge&lt;/strong&gt;&lt;/a&gt;&lt;strong style="vertical-align: baseline;"&gt; agent: &lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;A&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;ccelerates high-stakes legal and contracting work by combining institutional intelligence with a suite of agents that execute deep legal research, high-volume document analysis, and regulatory compliance screening.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Global systems integrators &amp;amp; tech partners:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Strategic partnerships with &lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;Accenture&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;, &lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;Deloitte&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;, &lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;Devoteam&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;, &lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;Factor Law, KPMG&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;, &lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;Tribe.ai&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;, &lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;Valtech&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;, &lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;Zazmic&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;, &lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;Zencore&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;, and &lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;66degrees&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; ensure legal teams can customize, integrate, and scale these capabilities across complex enterprise architectures without vendor lock-in.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;&lt;/div&gt;
&lt;div class="block-video"&gt;



&lt;div class="article-module article-video "&gt;
  &lt;figure&gt;
    &lt;a class="h-c-video h-c-video--marquee"
      href="https://youtube.com/watch?v=mMDDreSWhAg"
      data-glue-modal-trigger="uni-modal-mMDDreSWhAg-"
      data-glue-modal-disabled-on-mobile="true"&gt;

      
        

        &lt;div class="article-video__aspect-image"
          style="background-image: url(https://storage.googleapis.com/gweb-cloudblog-publish/images/Screenshot_2026-08-24_6.35.39_PM.max-1000x1000.png);"&gt;
          &lt;span class="h-u-visually-hidden"&gt;Gemini Enterprise for Legal&lt;/span&gt;
        &lt;/div&gt;
      
      &lt;svg role="img" class="h-c-video__play h-c-icon h-c-icon--color-white"&gt;
        &lt;use xlink:href="#mi-youtube-icon"&gt;&lt;/use&gt;
      &lt;/svg&gt;
    &lt;/a&gt;

    
  &lt;/figure&gt;
&lt;/div&gt;

&lt;div class="h-c-modal--video"
     data-glue-modal="uni-modal-mMDDreSWhAg-"
     data-glue-modal-close-label="Close Dialog"&gt;
   &lt;a class="glue-yt-video"
      data-glue-yt-video-autoplay="true"
      data-glue-yt-video-height="99%"
      data-glue-yt-video-vid="mMDDreSWhAg"
      data-glue-yt-video-width="100%"
      href="https://youtube.com/watch?v=mMDDreSWhAg"
      ng-cloak&gt;
   &lt;/a&gt;
&lt;/div&gt;

&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;p style="text-align: center;"&gt;&lt;em&gt;Gemini Enterprise for Legal offers leading firms a way to manage modern legal work with a secure agentic platform&lt;/em&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;Developed alongside leading global law firms&lt;/span&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;We are working closely with leading law firms, including Cleary Gottlieb, Freshfields, Weil, and Williams &amp;amp; Connolly, to ensure these capabilities address the realities of sophisticated legal practice.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="font-style: italic; vertical-align: baseline;"&gt;“Cleary is committed to embedding AI into our workflows in strategic and competitive ways. Using Google’s Gemini Enterprise, which can slot in seamlessly with other daily work tools, we can unlock greater efficiencies for our teams and help them deliver even higher quality work for our clients.” &lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;— &lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;Jeff Karpf&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;, Managing Partner, Cleary Gottlieb.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="font-style: italic; vertical-align: baseline;"&gt;“The legal sector is entering a period of accelerated change and transformation. For Freshfields the opportunity lies in how effectively we combine frontier technology like Gemini Enterprise for Legal with our expertise, robust governance and institutional knowledge to create value for our clients. Our strategic, multi-year partnership with Google Cloud is helping us accelerate that work and enhance how we deliver legal services.” &lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;— &lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;Alan Mason&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;, Global Managing Partner, Freshfields.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="font-style: italic; vertical-align: baseline;"&gt;“We’re thrilled to partner with Google Cloud in the early adoption of Gemini Enterprise for Legal. We look forward to integrating Google’s technology to streamline workflow and further support our litigators in shaping outcomes critical to our clients’ futures.” &lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;— &lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;Joe Petrosinelli&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;, Chairman, Williams &amp;amp; Connolly.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="font-style: italic; vertical-align: baseline;"&gt;"Our collaboration with Google gives us early access to emerging capabilities while allowing us to help shape the platform based on the realities of sophisticated legal practice. The result is technology that helps us continue to deliver the innovative, high-quality service our clients expect from Weil. We are looking forward to working with Google Cloud engineers and the Gemini Enterprise product team as we further innovate and evolve our AI capabilities." &lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;— &lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;Ramona Nee&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;, Incoming Executive Partner, Weil.&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-video"&gt;



&lt;div class="article-module article-video "&gt;
  &lt;figure&gt;
    &lt;a class="h-c-video h-c-video--marquee"
      href="https://youtube.com/watch?v=IUOYX4_0lUY"
      data-glue-modal-trigger="uni-modal-IUOYX4_0lUY-"
      data-glue-modal-disabled-on-mobile="true"&gt;

      
        

        &lt;div class="article-video__aspect-image"
          style="background-image: url(https://storage.googleapis.com/gweb-cloudblog-publish/images/Screenshot_2026-08-24_6.37.11_PM.max-1000x1000.png);"&gt;
          &lt;span class="h-u-visually-hidden"&gt;Weil Scales AI-Driven Judicial Insights With Gemini Enterprise&lt;/span&gt;
        &lt;/div&gt;
      
      &lt;svg role="img" class="h-c-video__play h-c-icon h-c-icon--color-white"&gt;
        &lt;use xlink:href="#mi-youtube-icon"&gt;&lt;/use&gt;
      &lt;/svg&gt;
    &lt;/a&gt;

    
  &lt;/figure&gt;
&lt;/div&gt;

&lt;div class="h-c-modal--video"
     data-glue-modal="uni-modal-IUOYX4_0lUY-"
     data-glue-modal-close-label="Close Dialog"&gt;
   &lt;a class="glue-yt-video"
      data-glue-yt-video-autoplay="true"
      data-glue-yt-video-height="99%"
      data-glue-yt-video-vid="IUOYX4_0lUY"
      data-glue-yt-video-width="100%"
      href="https://youtube.com/watch?v=IUOYX4_0lUY"
      ng-cloak&gt;
   &lt;/a&gt;
&lt;/div&gt;

&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;p style="text-align: center;"&gt;&lt;em&gt;Weil scales judicial insights with Benchmark, built on Gemini Enterprise&lt;/em&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;Built on an enterprise-grade foundation&lt;/span&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Confidentiality is not a feature of legal technology; it is the precondition for using any at all. Because Gemini Enterprise for Legal runs on Google Cloud infrastructure and the Gemini Enterprise platform, the permissions and access controls your firm already maintains are the boundaries the platform operates within — not settings it asks you to reconstruct.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Client data, firm-specific playbooks, intellectual property, custom agents, and model outputs stay private to your organization, and are never used to train or fine-tune Google's foundation models. Because we operate the full stack, from infrastructure and models through the application layer, we can tune performance and cost together — so expanding what your teams can take on does not mean expanding spend at the same rate.&lt;/span&gt;&lt;/p&gt;
&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;This is just the beginning&lt;/span&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;The launch of Gemini Enterprise for Legal represents another defining step in delivering on the promise of Gemini Enterprise: bringing the best of Google AI to every professional, for every workflow, natively tailored to the way they work.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;a href="https://cloud.google.com/ai/legal"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Gemini Enterprise for Legal&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; is available in &lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;preview&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; today, launching alongside our new purpose-built solution for &lt;/span&gt;&lt;a href="https://cloud.google.com/ai/financial-services"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Financial Services&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;. We invite law firms and legal teams to explore how &lt;/span&gt;&lt;a href="https://cloud.google.com/gemini-enterprise"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Gemini Enterprise&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; can transform their most critical workflows, with solutions for Healthcare, Life Sciences, and other Professional Services on the horizon. &lt;/span&gt;&lt;/p&gt;&lt;/div&gt;</description><pubDate>Tue, 25 Aug 2026 12:00:00 +0000</pubDate><guid>https://cloud.google.com/blog/products/ai-machine-learning/introducing-gemini-enterprise-for-legal/</guid><category>AI &amp; Machine Learning</category><media:content height="540" url="https://storage.googleapis.com/gweb-cloudblog-publish/images/Gemini_Enterprise_for_legal.max-600x600.jpg" width="540"></media:content><og xmlns:og="http://ogp.me/ns#"><type>article</type><title>Now introducing Gemini Enterprise for Legal</title><description></description><image>https://storage.googleapis.com/gweb-cloudblog-publish/images/Gemini_Enterprise_for_legal.max-600x600.jpg</image><site_name>Google</site_name><url>https://cloud.google.com/blog/products/ai-machine-learning/introducing-gemini-enterprise-for-legal/</url></og><author xmlns:author="http://www.w3.org/2005/Atom"><name>Thomas Kurian</name><title>CEO, Google Cloud</title><department></department><company></company></author></item><item><title>Now introducing Gemini Enterprise for Financial Services</title><link>https://cloud.google.com/blog/products/ai-machine-learning/introducing-gemini-enterprise-for-financial-services/</link><description>&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Protecting capital in today's markets requires immense speed and precision. A financial analyst preparing a deal memo works across licensed market data, internal models, and confidential client files. General-purpose AI lacks the real-time accuracy, verifiable data lineage, and strict security that financial institutions demand. While model intelligence is necessary, without deep integration into trusted financial systems, it is not sufficient.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Making AI genuinely useful inside an industry requires four things, together: &lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;domain expertise encoded into reusable skills, secure connections to the systems and data the work depends on, agents that can act inside real workflows, and an open ecosystem that extends and scales all of it &lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;— with governance running underneath all four.&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt; Each is valuable alone. Only together do they produce something an institution can actually put into production and see true return on investment.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Today, we are delivering on this vision with &lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;Gemini Enterprise for Financial Services&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;, bringing Google’s agentic AI directly into the workflows of capital markets and corporate banking.&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-video"&gt;



&lt;div class="article-module article-video "&gt;
  &lt;figure&gt;
    &lt;a class="h-c-video h-c-video--marquee"
      href="https://youtube.com/watch?v=dnwmxz2vkQQ"
      data-glue-modal-trigger="uni-modal-dnwmxz2vkQQ-"
      data-glue-modal-disabled-on-mobile="true"&gt;

      
        

        &lt;div class="article-video__aspect-image"
          style="background-image: url(https://storage.googleapis.com/gweb-cloudblog-publish/images/image2_6AhQtL3.max-1000x1000.png);"&gt;
          &lt;span class="h-u-visually-hidden"&gt;Gemini Enterprise for Financial Services&lt;/span&gt;
        &lt;/div&gt;
      
      &lt;svg role="img" class="h-c-video__play h-c-icon h-c-icon--color-white"&gt;
        &lt;use xlink:href="#mi-youtube-icon"&gt;&lt;/use&gt;
      &lt;/svg&gt;
    &lt;/a&gt;

    
  &lt;/figure&gt;
&lt;/div&gt;

&lt;div class="h-c-modal--video"
     data-glue-modal="uni-modal-dnwmxz2vkQQ-"
     data-glue-modal-close-label="Close Dialog"&gt;
   &lt;a class="glue-yt-video"
      data-glue-yt-video-autoplay="true"
      data-glue-yt-video-height="99%"
      data-glue-yt-video-vid="dnwmxz2vkQQ"
      data-glue-yt-video-width="100%"
      href="https://youtube.com/watch?v=dnwmxz2vkQQ"
      ng-cloak&gt;
   &lt;/a&gt;
&lt;/div&gt;

&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;p style="text-align: center;"&gt;&lt;em&gt;Bringing Gemini Enterprise into the workflows of capital markets and corporate banking&lt;/em&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;Four components, built for financial work &lt;/span&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Gemini Enterprise for Financial Services delivers an integrated, secure environment configured for rapid deployment with four core components:&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;1. Purpose-built financial skills.&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Skills are reusable packages of instructions and context that teach an agent to run a specialized task the way your institution runs it — applying custom formatting to a report, pulling a specific data cut, following a defined research methodology. They are available inside the Financial Research agent and to any agent your teams build.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;2. Secure Model Context Protocol (MCP) connectors.&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Direct integrations, using MCP, into essential financial platforms and licensed data sources, configured inside your own environment. Access stays bound by the entitlements you already maintain — licensed data stays licensed, and permissioned data stays permissioned.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;3. Agents that act. &lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;At its core is the Financial Research agent which is a Google-built, Google-managed agent that runs end-to-end research with full explainability. It ships with more than 50 foundational skills and exposes its reasoning through confidence scores, explicit methodologies, data snapshots for auditing, and precise source citations. Analysts can use it directly in the Gemini Enterprise app or wire it into existing agent workflows through Agent-to-Agent (A2A) APIs, and it connects to enterprise data sources over MCP to produce reports and documents in the formats your teams already use. Alongside it, out-of-the-box partner agents cover other workflows and extend the capabilities further.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;4. &lt;strong style="vertical-align: baseline;"&gt;An open partner ecosystem.&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; &lt;/span&gt;&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;&lt;span style="vertical-align: baseline;"&gt;Scale with global systems integrators and specialized fintech providers including &lt;/span&gt;&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;&lt;span style="vertical-align: baseline;"&gt;66degrees, Accenture, Artefact, Capgemini, Cognizant, Deloitte, Genpact, GFT Technologies, Infosys, KPMG, PwC, Quantiphi, Slalom, Tribe AI, and Zencore &lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;to customize and integrate the platform into your own architecture, without vendor lock-in.&lt;/span&gt;&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;Running underneath: a governed control plane.&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; A single dashboard for IT and risk teams that natively enforces security policies (VPC, CMEK), maintains private data isolation, and holds every output to verifiable grounding with traceable citations.&lt;/span&gt;&lt;/p&gt;
&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;Unlocking high-value workflows with domain-specific skills&lt;/span&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Whether used by private equity specialists, wealth managers, or compliance teams, the solution adapts to diverse workflows like credit risk assessment, portfolio monitoring, market news synthesis, and investigative financial research:&lt;/span&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Elevate advisor insights:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Equips relationship managers and advisors with AI-generated insights, personalized recommendations, and tailored artifacts, enabling higher-quality conversations and fostering loyalty. &lt;/span&gt;&lt;/li&gt;
&lt;li role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Deepen Know Your Customer (KYC) research and analysis: &lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;Modernizes onboarding and Know Your Customer (KYC) workflows across private banking and prime brokerage by using multi-format ingestion (PDFs, Excel, SEC filings) to map complex corporate hierarchies, evaluate risk personas, and resolve ultimate beneficial owners (UBOs).&lt;/span&gt;&lt;/li&gt;
&lt;li role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Enhance portfolio resilience:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Helps tra&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;ding desks deal with sudden macroeconomic shocks. It reduces complex bond portfolio risk exposure analysis to a sub-5-minute execution, complete with automated duration-hedging strategy suggestions.&lt;/span&gt;&lt;/li&gt;
&lt;li role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Uncover credit market opportunities:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Transforms credit data into actionable trade ideas by identifying and isolating potential mispricings. This enables teams to expand trading volumes while lowering back-office risk and underwriting latency.&lt;/span&gt;&lt;/li&gt;
&lt;li role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Accelerate bond issuance: &lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;Compresses&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt; client pitch presentation timelines from days to minutes &lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;so that fixed-income and underwriting teams &lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;can proactively target prospects, increase deal capacity, and secure a crucial first-mover advantage to help win more business.&lt;/span&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;Open ecosystem of connectors across the financial technology stack&lt;/span&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Gemini Enterprise connects directly to core financial systems via secure &lt;a href="https://cloud.google.com/gemini-enterprise/connectors?e=48754805"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;MCP connectors&lt;/span&gt;&lt;/a&gt;. Access is bound by existing role-based controls, ensuring verifiable grounding and precise source citations:&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;Productivity and collaboration:&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Google Workspace:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Enables seamless analysis and live artifact generation across Docs, Sheets, and Slides while adhering to enterprise DLP policies.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Microsoft 365:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Integrates directly with Excel, Word, and PowerPoint to populate financial models, research memos, and client pitch decks.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;Market data and financial fundamentals:&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Daloopa:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Provides the structured, source-linked financial data layer that enables finance professionals and AI tools to produce accurate and auditable results. &lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;&lt;strong style="vertical-align: baseline;"&gt;FactSet:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Enables secure, authorized access to FactSet's multi-asset class financial and non-financial datasets, powering reliable AI-driven workflows with fully auditable, compliant insights. &lt;/span&gt;&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Finnhub:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Provides real-time financial APIs, global fundamentals, and earnings call transcripts for in-depth financial research.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Fiscal.ai:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; &lt;span style="vertical-align: baseline;"&gt;Delivers institutional-grade financial data within minutes of earnings, covering financials, news, ownership, segments &amp;amp; KPIs, filings, and earnings call transcripts. &lt;/span&gt;&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Guidepoint:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Connects to primary research insights and expert network transcripts to inform and validate investment theses.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;span style="vertical-align: baseline;"&gt;&lt;strong style="vertical-align: baseline;"&gt;LSEG:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Provides access to a broad range of trusted financial data, analytical models, indices and news. Enabling customers turnkey access to trusted financial intelligence across every stage of the investment lifecycle.&lt;/span&gt;&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;&lt;strong style="vertical-align: baseline;"&gt;S&amp;amp;P Global:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Integrates cited, verifiable S&amp;amp;P Global data for a range of workflows, from financial analysis, to peer benchmarking, industry research, and more.&lt;/span&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;Risk, ratings, and private markets:&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Moody’s:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Brings ratings, default risk models, and real time news fused into one lens for counterparty risk assessment. &lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;MSCI:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Connects to proprietary indexes, data and models spanning public and private assets and also provides risk analytics and factor exposures. &lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;PitchBook:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Provides comprehensive data and research on private equity, venture capital, credit, M&amp;amp;A, and public markets, including, company financials, deal terms, valuations, and fund performance.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;Regulatory and corporate records:&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;SEC Edgar:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Delivers instant, verifiable retrieval of statutory filings, 10-Ks, 10-Qs, and 8-Ks with precise citation mapping.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Dun &amp;amp; Bradstreet:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Accelerates commercial onboarding and KYB verification through direct access to global corporate hierarchy records.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;Digital assets and indices:&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;CoinDesk Data and Indices:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Supplies institutional-grade digital asset pricing, benchmark indices, and crypto market intelligence for multi-asset strategies.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;&lt;/div&gt;
&lt;div class="block-video"&gt;



&lt;div class="article-module article-video "&gt;
  &lt;figure&gt;
    &lt;a class="h-c-video h-c-video--marquee"
      href="https://youtube.com/watch?v=GZCSuRKDfBc"
      data-glue-modal-trigger="uni-modal-GZCSuRKDfBc-"
      data-glue-modal-disabled-on-mobile="true"&gt;

      
        

        &lt;div class="article-video__aspect-image"
          style="background-image: url(https://storage.googleapis.com/gweb-cloudblog-publish/images/Screenshot_2026-08-24_6.25.38_PM.max-1000x1000.png);"&gt;
          &lt;span class="h-u-visually-hidden"&gt;Introducing Gemini Enterprise for Financial Services&lt;/span&gt;
        &lt;/div&gt;
      
      &lt;svg role="img" class="h-c-video__play h-c-icon h-c-icon--color-white"&gt;
        &lt;use xlink:href="#mi-youtube-icon"&gt;&lt;/use&gt;
      &lt;/svg&gt;
    &lt;/a&gt;

    
  &lt;/figure&gt;
&lt;/div&gt;

&lt;div class="h-c-modal--video"
     data-glue-modal="uni-modal-GZCSuRKDfBc-"
     data-glue-modal-close-label="Close Dialog"&gt;
   &lt;a class="glue-yt-video"
      data-glue-yt-video-autoplay="true"
      data-glue-yt-video-height="99%"
      data-glue-yt-video-vid="GZCSuRKDfBc"
      data-glue-yt-video-width="100%"
      href="https://youtube.com/watch?v=GZCSuRKDfBc"
      ng-cloak&gt;
   &lt;/a&gt;
&lt;/div&gt;

&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;p style="text-align: center;"&gt;&lt;em&gt;&lt;span style="vertical-align: baseline;"&gt;Introducing Gemini Enterprise for Financial Services&lt;/span&gt;&lt;/em&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;Third-party agents and implementation partners&lt;/span&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Organizations can deploy out-of-the-box partner agents or collaborate with global systems integrators to scale custom capabilities without vendor lock-in:&lt;/span&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;a href="https://console.cloud.google.com/marketplace/product/prod-dnb-mp-saas-publicae85678/business-verification-a2a" rel="noopener" target="_blank"&gt;&lt;strong style="text-decoration: underline; vertical-align: baseline;"&gt;D&amp;amp;B Business Verification&lt;/strong&gt;&lt;/a&gt;&lt;strong style="vertical-align: baseline;"&gt; agent:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Accelerates commercial onboarding and strengthens KYC compliance.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;FlowX agents:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; &lt;span style="vertical-align: baseline;"&gt;Automate loan pack completeness check, document reconciliation and many other mission critical processes for financial institutions&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;.&lt;/span&gt;&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;a href="https://console.cloud.google.com/marketplace/product/obin-public/obin-financial-agent"&gt;&lt;strong style="text-decoration: underline; vertical-align: baseline;"&gt;Obin Financial&lt;/strong&gt;&lt;/a&gt;&lt;strong style="vertical-align: baseline;"&gt; agent:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; &lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;Helps asset management, commercial lending, and insurance teams accelerate complex financial analyses.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;a href="https://console.cloud.google.com/marketplace/product/kensho-groundings-poc/sp-global-data-retrieval-agent"&gt;&lt;strong style="text-decoration: underline; vertical-align: baseline;"&gt;S&amp;amp;P Global&lt;/strong&gt;&lt;/a&gt;&lt;strong style="vertical-align: baseline;"&gt; agents:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; &lt;a href="https://console.cloud.google.com/marketplace/product/kensho-groundings-poc/sp-global-data-retrieval-agent"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Data Retrieval Agent&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; for multi-step analysis, report generation, research workflows, and the &lt;/span&gt;&lt;a href="https://console.cloud.google.com/marketplace/product/p-s1-marketplace/spgi-hzna-tfa"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Horizons Agents&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; that help turn complex energy and sustainability data into fast insights for finance workflows.&lt;/span&gt;&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Global systems integrators and tech partners: &lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;Strategic partnerships connect firms with specialist FinTech and leading global systems integrators, including 66degrees, Accenture, Artefact, Capgemini, Cognizant, Deloitte, Genpact, GFT Technologies, Infosys, KPMG, PwC, Quantiphi, Slalom, Tribe AI, and Zencore to manage custom configurations and deploy specialized capabilities at a global scale.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;Developed alongside leading global financial institutions&lt;/span&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;We are developing these capabilities in close collaboration with financial institutions, including &lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;Deutsche Bank and CME Group&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;, to ensure they reflect the operational realities of the industry.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;“As a design partner for the Financial Research agent, Deutsche Bank has helped shape this capability in view of the realities of a highly regulated industry – from data protection and governance to the workflows our teams use every day,” said Marie-Jeanne Deverdun, Chief Technology, Data and Innovation Officer, and Member of the Deutsche Bank Management Board. “Starting in the Corporate Bank, we see significant potential to reduce manual research effort, improve the consistency and auditability of outputs, and give our teams more time for client conversations. This is an important step in applying AI where it can make a practical difference: safely, responsibly and at scale.”&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;This launch builds on the rapidly growing momentum of Gemini Enterprise, with many leading financial institutions like &lt;/span&gt;&lt;a href="https://www.googlecloudpresscorner.com/2025-12-08-BNY-Collaborates-with-Google-Cloud-to-Advance-its-Eliza-AI-Platform-with-Gemini-Enterprise" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;BNY&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;,&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt; &lt;/span&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt; &lt;/span&gt;&lt;a href="https://www.citigroup.com/global/news/press-release/2026/citi-wealth-unveils-citi-sky-ai-powered-member-google-cloud-deepmind-technologies" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Citi Wealth&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;,&lt;/span&gt;&lt;a href="https://cloud.google.com/customers/lloydsbankinggroup?e=48754805"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt; Lloyds Banking Group,&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; &lt;/span&gt;&lt;a href="https://www.googlecloudpresscorner.com/2025-10-09-Macquarie-Bank-Democratizes-Agentic-AI,-Scaling-Customer-Innovation-with-Gemini-Enterprise" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Macquarie Bank&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;, and &lt;/span&gt;&lt;a href="https://www.googlecloudpresscorner.com/2025-10-09-SIGNAL-IDUNA-Rolls-Out-Gemini-Enterprise-for-over-10,000-Employees" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Signal Iduna&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; using it &lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;to equip their workforce with advanced, agentic workflow tools to drive growth and efficiency. &lt;/span&gt;&lt;/p&gt;
&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;Built on an enterprise-grade foundation&lt;/span&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Because Gemini Enterprise for Financial Services runs on Google Cloud infrastructure and the Gemini Enterprise platform, organizations get the security, governance, compliance, and cost-management capabilities they expect from an enterprise platform. &lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Customer data, business rules, intellectual property, custom agents, and model outputs remain private to their organization. Your data is never used to train or fine-tune Google’s foundation models. Furthermore, our full-stack approach - from infrastructure and models to the application layer - allows us to optimize performance and cost, helping organizations maximize the value of their AI investments. &lt;/span&gt;&lt;/p&gt;
&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;This is just the beginning&lt;/span&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;&lt;a href="https://cloud.google.com/ai/financial-services"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Gemini Enterprise for Financial Services&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; is available in &lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;preview&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; today, launching alongside our new purpose-built solution for &lt;/span&gt;&lt;a href="https://cloud.google.com/ai/legal"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Legal&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;. We invite enterprise leaders in financial institutions to explore how &lt;/span&gt;&lt;a href="https://cloud.google.com/gemini-enterprise"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Gemini Enterprise&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; can transform their most critical workflows, with solutions for Healthcare, Life Sciences, and other Professional Services on the horizon. &lt;/span&gt;&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;</description><pubDate>Tue, 25 Aug 2026 12:00:00 +0000</pubDate><guid>https://cloud.google.com/blog/products/ai-machine-learning/introducing-gemini-enterprise-for-financial-services/</guid><category>AI &amp; Machine Learning</category><media:content height="540" url="https://storage.googleapis.com/gweb-cloudblog-publish/images/Gemini_Enterprise_for_Finserve_ru2ruLE.max-600x600.jpg" width="540"></media:content><og xmlns:og="http://ogp.me/ns#"><type>article</type><title>Now introducing Gemini Enterprise for Financial Services</title><description></description><image>https://storage.googleapis.com/gweb-cloudblog-publish/images/Gemini_Enterprise_for_Finserve_ru2ruLE.max-600x600.jpg</image><site_name>Google</site_name><url>https://cloud.google.com/blog/products/ai-machine-learning/introducing-gemini-enterprise-for-financial-services/</url></og><author xmlns:author="http://www.w3.org/2005/Atom"><name>Thomas Kurian</name><title>CEO, Google Cloud</title><department></department><company></company></author></item><item><title>Cloud CISO Perspectives: Sticking to security fundamentals in the AI era</title><link>https://cloud.google.com/blog/products/identity-security/cloud-ciso-perspectives-sticking-to-security-fundamentals-in-the-ai-era/</link><description>&lt;div class="block-paragraph"&gt;&lt;p data-block-key="eucpw"&gt;Welcome to the first Cloud CISO Perspectives for August 2026. Today, Chris Betz explains why the AI era makes it more important than ever to lean into security fundamentals.&lt;/p&gt;&lt;p data-block-key="3e9dh"&gt;As with all Cloud CISO Perspectives, the contents of this newsletter are posted to the &lt;a href="https://cloud.google.com/blog/products/identity-security/"&gt;Google Cloud blog&lt;/a&gt;. If you’re reading this on the website and you’d like to receive the email version, you can &lt;a href="https://cloud.google.com/resources/google-cloud-ciso-newsletter-signup"&gt;subscribe here&lt;/a&gt;.&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-aside"&gt;&lt;dl&gt;
    &lt;dt&gt;aside_block&lt;/dt&gt;
    &lt;dd&gt;&amp;lt;ListValue: [StructValue([(&amp;#x27;title&amp;#x27;, &amp;#x27;Get vital board insights with Google Cloud&amp;#x27;), (&amp;#x27;body&amp;#x27;, &amp;lt;wagtail.rich_text.RichText object at 0x7f1876eb80a0&amp;gt;), (&amp;#x27;btn_text&amp;#x27;, &amp;#x27;Visit the hub&amp;#x27;), (&amp;#x27;href&amp;#x27;, &amp;#x27;https://cloud.google.com/solutions/security/board-of-directors?utm_source=cgc-site&amp;amp;utm_medium=et&amp;amp;utm_campaign=FY26-Q2-GLOBAL-GCP39634-email-dl-dgcsm-CISOP-NL-177159&amp;amp;utm_content=-&amp;amp;utm_term=-&amp;#x27;), (&amp;#x27;image&amp;#x27;, &amp;lt;GAEImage: GCAT-replacement-logo-A&amp;gt;)])]&amp;gt;&lt;/dd&gt;
&lt;/dl&gt;&lt;/div&gt;
&lt;div class="block-paragraph"&gt;&lt;h3 data-block-key="hswvv"&gt;&lt;b&gt;How to stay strong with security fundamentals in the AI era&lt;/b&gt;&lt;/h3&gt;&lt;p data-block-key="eoh1k"&gt;&lt;i&gt;By Chris Betz, CISO, Google Cloud&lt;/i&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-paragraph_with_image"&gt;&lt;div class="article-module h-c-page"&gt;
  &lt;div class="h-c-grid uni-paragraph-wrap"&gt;
    &lt;div class="uni-paragraph
      h-c-grid__col h-c-grid__col--8 h-c-grid__col-m--6 h-c-grid__col-l--6
      h-c-grid__col--offset-2 h-c-grid__col-m--offset-3 h-c-grid__col-l--offset-3"&gt;

      






  

    &lt;figure class="article-image--wrap-small
      
      "
      &gt;

      
      
        
        &lt;img
            src="https://storage.googleapis.com/gweb-cloudblog-publish/images/Chris_Betz_Google-9779.max-1000x1000.jpg"
        
          alt="Chris Betz Google-9779"&gt;
        
        &lt;/a&gt;
      
        &lt;figcaption class="article-image__caption "&gt;&lt;p data-block-key="nj7d4"&gt;Chris Betz, CISO, Google Cloud&lt;/p&gt;&lt;/figcaption&gt;
      
    &lt;/figure&gt;

  





      &lt;p data-block-key="0jyqm"&gt;As AI accelerates the capabilities of adversaries, foundational strength becomes the primary differentiator between resilience and vulnerability. It’s a dangerous and unfortunately common misconception that traditional security fundamentals are becoming obsolete. For CISOs, the challenge is to adopt new AI technology securely while scaling essential, effective defensive practices to move at the speed of the adversary.&lt;/p&gt;&lt;p data-block-key="anha5"&gt;For both attackers and defenders, AI has been a catalyst for optimization and innovation. While traditional automation has allowed us to perform repetitive tasks at scale, AI enables both sides to execute highly-specific, customized actions at massive scale and unprecedented speed.&lt;/p&gt;
    &lt;/div&gt;
  &lt;/div&gt;
&lt;/div&gt;

&lt;/div&gt;
&lt;div class="block-pull_quote"&gt;&lt;div class="uni-pull-quote h-c-page"&gt;
  &lt;section class="h-c-grid"&gt;
    &lt;div class="uni-pull-quote__wrapper h-c-grid__col h-c-grid__col--8 h-c-grid__col-m--6 h-c-grid__col-l--6
      h-c-grid__col--offset-2 h-c-grid__col-m--offset-3 h-c-grid__col-l--offset-3"&gt;
      &lt;div class="uni-pull-quote__inner-wrapper h-c-copy h-c-copy"&gt;
        &lt;q class="uni-pull-quote__text"&gt;Collectively, these technologies reduce the attack surface and contribute to the deep context that defensive AI needs to be a business enabler — and create the necessary conditions for successful AI-powered defenses.&lt;/q&gt;

        
      &lt;/div&gt;
    &lt;/div&gt;
  &lt;/section&gt;
&lt;/div&gt;

&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;We can see the threat developing almost in real-time. Adversaries are &lt;/span&gt;&lt;a href="https://cloud.google.com/security/resources/ai-risk-and-resilience"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;deploying new malware&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; with just-in-time AI that dynamically generates malicious scripts and obfuscates code mid-execution to evade detection. They use sophisticated vishing and deepfakes for identity theft and business email compromise. We even see unauthorized AI tools lead to the rise of shadow agents. &lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Defending against AI powered security threats requires more than accelerating current security practices; it means stepping back and beginning with the security foundation and layered defenses. It’s critically important to build and use a layered defense with the right guardrails — foundational cybersecurity building blocks that we’ve been investing in for years.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Doubling down on this foundation: technologies like multi-factor authentication (MFA), &lt;/span&gt;&lt;a href="https://blog.google/security/going-beyond-zero-a-new-paradigm-for-enterprise-security/" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Zero Trust frameworks&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;, consistent system patching, and comprehensive detection and response. Collectively, these technologies reduce the attack surface and contribute to the &lt;/span&gt;&lt;a href="https://cloud.google.com/blog/products/identity-security/cloud-ciso-perspectives-how-ai-leverages-deep-context-defenders-advantage"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;deep context that defensive AI needs&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; to be a business enabler — and create the necessary conditions for successful AI-powered defenses.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;Revolutionizing vulnerability management&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;In just a few short years, identifying and fixing vulnerabilities has evolved from a mostly laborious, manual process to one driven by AI tools discovering vulnerabilities at volumes never seen before. Further, the time to exploit window has essentially been eliminated.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;However, it’s not enough to merely discover vulnerabilities, especially at today’s volumes. You still need to prioritize fixing those that have the most critical impact on your systems and networks first, and that necessitates an equally-rapid response in smart mitigation. &lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Organizations use multiple models to scan for flaws and then suggest high-quality code fixes that engineers can quickly move into production, leveraging capabilities like &lt;/span&gt;&lt;a href="https://cloud.google.com/security/ai-threat-defense"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;AI Threat Defense&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;. AI allows us to automate the entire software development lifecycle, from discovery to testing and deployment, ensuring that our defensive posture evolves faster than the threats targeting us.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;Enhancing threat modeling&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;We’re also seeing the fundamental concept of threat modeling have an outsized impact. Doing threat modeling well requires bringing context together from your code, your cloud architecture, system design, and network pathways. &lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;While it isn’t easy, using AI can scale our ability to bring that data together into a coherent picture. Teams have been experimenting with multi-AI models to collect system information and enumerate threats. &lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;As I noted in June, &lt;/span&gt;&lt;a href="https://cloud.google.com/blog/products/identity-security/cloud-ciso-perspectives-how-google-cloud-security-uses-ai-internally/"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;engineering teams at Google Cloud&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; now route product launches through an agent-based security review pipeline. High-risk indicators automatically get flagged for human review, while we’ve replaced static threat models with dynamic product dossiers that update in real-time.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;The CISO as a strategic business leader&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;The most effective security leaders that I know today are more than just technologists: They are strategic business leaders. The intense global focus on AI vulnerabilities has brought cybersecurity to the &lt;/span&gt;&lt;a href="https://cloud.google.com/blog/products/identity-security/cloud-ciso-perspectives-why-ai-threat-defense-is-the-new-boardroom-baseline"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;forefront of boardroom and executive attention&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; like never before.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;This visibility is an opportunity to lead. We CISOs are expected to communicate with clarity, from the board to the C-suite to the security teams who look to them on a daily basis, demonstrating their ability as capable strategists who can navigate the complexities of AI while safeguarding the organization's growth. &lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;By aligning security fundamentals with business objectives and using AI to enhance defense, we can lead our organizations securely into the future.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;To learn more about building and maintaining strong security foundations in the AI era, read our newest &lt;/span&gt;&lt;a href="https://cloud.google.com/security/resources/cyber-snapshot-reports"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Defender’s Advantage: Cyber Snapshot Report&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;.&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-aside"&gt;&lt;dl&gt;
    &lt;dt&gt;aside_block&lt;/dt&gt;
    &lt;dd&gt;&amp;lt;ListValue: [StructValue([(&amp;#x27;title&amp;#x27;, &amp;#x27;Learn something new&amp;#x27;), (&amp;#x27;body&amp;#x27;, &amp;lt;wagtail.rich_text.RichText object at 0x7f1876eb8d00&amp;gt;), (&amp;#x27;btn_text&amp;#x27;, &amp;#x27;Watch now&amp;#x27;), (&amp;#x27;href&amp;#x27;, &amp;#x27;https://www.youtube.com/watch?v=CmGWIwgHR60&amp;#x27;), (&amp;#x27;image&amp;#x27;, &amp;lt;GAEImage: Cloud-CISO-Perspectives-logo-A&amp;gt;)])]&amp;gt;&lt;/dd&gt;
&lt;/dl&gt;&lt;/div&gt;
&lt;div class="block-paragraph"&gt;&lt;h3 data-block-key="4bd61"&gt;&lt;b&gt;In case you missed it&lt;/b&gt;&lt;/h3&gt;&lt;p data-block-key="5tvtn"&gt;Here are the latest updates, products, services, and resources from our security teams so far this month:&lt;/p&gt;&lt;ul&gt;&lt;li data-block-key="4h3h5"&gt;&lt;b&gt;Driving AI threat readiness with Wiz&lt;/b&gt;: Announcing new Wiz capabilities that can help organizations prepare for the AI era by expanding visibility and accelerating response, so your security teams can defend at machine speed. &lt;a href="https://www.wiz.io/blog/wiz-at-black-hat-2026" target="_blank"&gt;&lt;b&gt;Read more&lt;/b&gt;&lt;/a&gt;.&lt;/li&gt;&lt;li data-block-key="d5vd8"&gt;&lt;b&gt;PQC in Plaintext: Google Cloud’s post-quantum cryptography roadmap&lt;/b&gt;: We’ve long been actively working on and rolling out post-quantum cryptography in our infrastructure. Here’s our updated Google Cloud roadmap to migrate to PQC by 2029. &lt;a href="https://cloud.google.com/blog/products/identity-security/pqc-in-plaintext-google-clouds-post-quantum-cryptography-roadmap"&gt;&lt;b&gt;Read more&lt;/b&gt;&lt;/a&gt;.&lt;/li&gt;&lt;li data-block-key="6anll"&gt;&lt;b&gt;How Google Cloud detects, contains, and protects against emerging threats&lt;/b&gt;: Learn more about how Google Cloud empowers you with the tools, governance, and infrastructure you need to securely deploy workloads and maintain long-term trust. &lt;a href="https://cloud.google.com/blog/products/identity-security/how-google-cloud-detects-contains-and-protects-against-emerging-threats"&gt;&lt;b&gt;Read more&lt;/b&gt;&lt;/a&gt;.&lt;/li&gt;&lt;li data-block-key="9pqs5"&gt;&lt;b&gt;Privacy-first medical AI with MedPerf and Google Cloud&lt;/b&gt;: Discover how Google Cloud and MedPerf use Confidential Computing to enable secure, privacy-first collaborative medical AI evaluation. &lt;a href="https://cloud.google.com/blog/products/identity-security/privacy-first-medical-ai-with-medperf-and-google-cloud"&gt;&lt;b&gt;Read more&lt;/b&gt;&lt;/a&gt;.&lt;/li&gt;&lt;li data-block-key="djij2"&gt;&lt;b&gt;More cryptanalysis makes us all safer&lt;/b&gt;: Recent advances in frontier AI models do not signal the downfall of cryptography. Here’s why they’re best viewed as additional cryptanalysts. &lt;a href="https://bughunters.google.com/blog/more-cryptanalysis-makes-us-all-safer" target="_blank"&gt;&lt;b&gt;Read more&lt;/b&gt;&lt;/a&gt;.&lt;/li&gt;&lt;li data-block-key="dhi1p"&gt;&lt;b&gt;How layered defenses harden Chrome against abusive notifications&lt;/b&gt;: Learn how Chrome Security has collaborated with Firebase Cloud Messaging (FCM) and Safe Browsing to significantly reduce notification abuse, and improve the security and quality of the web ecosystem for everyone. &lt;a href="https://blog.google/security/the-multi-layered-defenses-that-harden-chrome-against-abusive-notifications/" target="_blank"&gt;&lt;b&gt;Read more&lt;/b&gt;&lt;/a&gt;.&lt;/li&gt;&lt;/ul&gt;&lt;p data-block-key="5nair"&gt;Please visit the Google Cloud blog for more security stories &lt;a href="https://cloud.google.com/blog/products/identity-security"&gt;published this month&lt;/a&gt;.&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-aside"&gt;&lt;dl&gt;
    &lt;dt&gt;aside_block&lt;/dt&gt;
    &lt;dd&gt;&amp;lt;ListValue: [StructValue([(&amp;#x27;title&amp;#x27;, &amp;#x27;Join the Google Cloud CISO Community&amp;#x27;), (&amp;#x27;body&amp;#x27;, &amp;lt;wagtail.rich_text.RichText object at 0x7f1876eb8160&amp;gt;), (&amp;#x27;btn_text&amp;#x27;, &amp;#x27;Learn more&amp;#x27;), (&amp;#x27;href&amp;#x27;, &amp;#x27;https://rsvp.withgoogle.com/events/google-cloud-ciso-community-interest-form-2026?utm_source=cgc-blog&amp;amp;utm_medium=blog&amp;amp;utm_campaign=FY25-Q1-global-GCP30328-physicalevent-er-dgcsm-parent-CISO-community-2025&amp;amp;utm_content=cisop_&amp;amp;utm_term=-&amp;#x27;), (&amp;#x27;image&amp;#x27;, &amp;lt;GAEImage: GCAT-replacement-logo-A&amp;gt;)])]&amp;gt;&lt;/dd&gt;
&lt;/dl&gt;&lt;/div&gt;
&lt;div class="block-paragraph"&gt;&lt;h3 data-block-key="29tyz"&gt;&lt;b&gt;Threat Intelligence news&lt;/b&gt;&lt;/h3&gt;&lt;ul&gt;&lt;li data-block-key="cm2sc"&gt;&lt;b&gt;Staying ahead of adversarial AI through agentic source code review&lt;/b&gt;: To help defenders implement agentic approaches similar to our approach at Google Cloud, we are sharing the details of our Agentic Vulnerability Discovery Harness architecture for the first time. AVDH can also be used alongside CodeMender’s ongoing scanning to create a two-layered defense strategy. &lt;a href="https://cloud.google.com/blog/topics/threat-intelligence/staying-ahead-of-adversarial-ai-through-agentic-source-code-review"&gt;&lt;b&gt;Read more&lt;/b&gt;&lt;/a&gt;.&lt;/li&gt;&lt;li data-block-key="6lp6o"&gt;&lt;b&gt;Cloud threat highlights from the first half of 2026&lt;/b&gt;: In the first half of 2026, Wiz's Research and CIRT teams tracked threats affecting thousands of cloud environments. We saw a notable increase in the volume of activity, with supply-chain attacks running at a previously unseen scale and developer toolchains and AI infrastructure drawing serious attention. &lt;a href="https://www.wiz.io/blog/cloud-threat-highlights-h1-2026" target="_blank"&gt;&lt;b&gt;Read more&lt;/b&gt;&lt;/a&gt;.&lt;/li&gt;&lt;li data-block-key="dtcu6"&gt;&lt;b&gt;Batten down your packages: Mitigation guidance for supply chain compromise&lt;/b&gt;: GTIG and Mandiant have tracked ongoing and increasing open source software supply chain compromise campaigns over the past several years. Here are our mitigation and hardening recommendations to secure software supply chains, including insights we have developed as a result of supporting customers. &lt;a href="https://cloud.google.com/blog/topics/threat-intelligence/mitigation-guidance-for-supply-chain-compromise"&gt;&lt;b&gt;Read more&lt;/b&gt;&lt;/a&gt;.&lt;/li&gt;&lt;li data-block-key="bqqfp"&gt;&lt;b&gt;Multi-brand vishing extortion targets financial services and enterprise cloud environments&lt;/b&gt;: Telemetry and infrastructure analysis reveal that UNC6671 has not disbanded. Instead, the threat group has diversified its operations across multiple extortion fronts including Redact, Pink, Helix, and Falcon and continues to rely on voice phishing to target enterprise employees. &lt;a href="https://cloud.google.com/blog/topics/threat-intelligence/unc6671-targets-financial-services-and-enterprise-cloud-environments"&gt;&lt;b&gt;Read more&lt;/b&gt;&lt;/a&gt;.&lt;/li&gt;&lt;li data-block-key="74ud9"&gt;&lt;b&gt;Keyv and cacheable npm package hijacked in supply chain attack&lt;/b&gt;: Wiz Research is actively investigating an ongoing software supply chain attack affecting multiple keyv/cacheable npm packages. &lt;a href="https://www.wiz.io/blog/keyv-and-cacheable-npm-supply-chain-attack" target="_blank"&gt;&lt;b&gt;Read more&lt;/b&gt;&lt;/a&gt;.&lt;/li&gt;&lt;li data-block-key="49ocf"&gt;&lt;b&gt;Inside the Metabase SQLi: Exploited in the wild&lt;/b&gt;: Wiz has reverse engineered Metabase CVE-2026-72898 with AI to accelerate defense. Here’s what we learned. &lt;a href="https://www.wiz.io/blog/inside-the-metabase-sqli-exploited-in-the-wild" target="_blank"&gt;&lt;b&gt;Read more&lt;/b&gt;&lt;/a&gt;.&lt;/li&gt;&lt;/ul&gt;&lt;p data-block-key="d9no2"&gt;Please visit the Google Cloud blog for more threat intelligence stories &lt;a href="https://cloud.google.com/blog/topics/threat-intelligence/"&gt;published this month&lt;/a&gt;.&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-paragraph"&gt;&lt;h3 data-block-key="rcfc5"&gt;&lt;b&gt;Now hear this: Podcasts from Google Cloud&lt;/b&gt;&lt;/h3&gt;&lt;ul&gt;&lt;li data-block-key="drbpp"&gt;&lt;b&gt;Cloud Security Podcast: All about Project Atlas, Wiz's AI vulnerability research&lt;/b&gt;: Near Orfeld, head of vulnerability research, Wiz, discusses how his team uses multi-agent AI systems for discovering high-impact zero-day vulnerabilities in cloud infrastructure. &lt;a href="https://www.youtube.com/watch?v=qRJJ9ekpuVg" target="_blank"&gt;&lt;b&gt;Listen here&lt;/b&gt;&lt;/a&gt;.&lt;/li&gt;&lt;li data-block-key="1chcd"&gt;&lt;b&gt;Cloud Security Podcast: How Google eliminates classes of vulnerabilities at scale&lt;/b&gt;: How do you build the foundations for a secure Google-scale enterprise that stays secure even if an AI is writing the code and nobody has time to review it? Christoph Kern, principal security engineer, Google, explores what secure-by-design really means in the AI era. &lt;a href="https://www.youtube.com/watch?v=43imRRfgLgc" target="_blank"&gt;&lt;b&gt;Listen here&lt;/b&gt;&lt;/a&gt;.&lt;/li&gt;&lt;/ul&gt;&lt;p data-block-key="fjjbt"&gt;To have our Cloud CISO Perspectives post delivered twice a month to your inbox, &lt;a href="https://cloud.google.com/resources/google-cloud-ciso-newsletter-signup"&gt;sign up for our newsletter&lt;/a&gt;. We’ll be back in a few weeks with more security-related updates from Google Cloud.&lt;/p&gt;&lt;/div&gt;</description><pubDate>Fri, 21 Aug 2026 16:00:00 +0000</pubDate><guid>https://cloud.google.com/blog/products/identity-security/cloud-ciso-perspectives-sticking-to-security-fundamentals-in-the-ai-era/</guid><category>Cloud CISO</category><category>AI &amp; Machine Learning</category><category>Security &amp; Identity</category><media:content height="540" url="https://storage.googleapis.com/gweb-cloudblog-publish/images/Cloud_CISO_Perspectives_header_4_Blue.max-600x600.png" width="540"></media:content><og xmlns:og="http://ogp.me/ns#"><type>article</type><title>Cloud CISO Perspectives: Sticking to security fundamentals in the AI era</title><description></description><image>https://storage.googleapis.com/gweb-cloudblog-publish/images/Cloud_CISO_Perspectives_header_4_Blue.max-600x600.png</image><site_name>Google</site_name><url>https://cloud.google.com/blog/products/identity-security/cloud-ciso-perspectives-sticking-to-security-fundamentals-in-the-ai-era/</url></og><author xmlns:author="http://www.w3.org/2005/Atom"><name>Chris Betz</name><title>CISO, Google Cloud</title><department></department><company></company></author></item><item><title>How agents can delegate better</title><link>https://cloud.google.com/blog/products/ai-machine-learning/how-agents-can-delegate-better/</link><description>&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;In any organizational behavior class, students will learn that effective delegation is among the most important skills for a seasoned leader. Getting meaningful work done involves careful coordination, starting with a subdivision of projects into manageable tasks, mapped onto the skills of the team, and assigned to the right people.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;At Google Cloud, we’re learning a similar lesson when it comes to building and deploying AI agents in enterprise workflows. These workflows are best approached by multi-agent systems that can break apart and execute complex tasks. To do so, AI agents need to become good delegators. &lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;To learn how, we turned to research from Google DeepMind. In their recent study titled &lt;/span&gt;&lt;a href="https://arxiv.org/abs/2602.11865" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Intelligent AI Delegation&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;, they prove how delegation itself involves intelligence: adaptive negotiations, aligning on formal contracts, and security guardrails. &lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;This work opens up new opportunities for customers building AI agents that can communicate, share tasks, and coordinate towards set objectives. Today, we’ll share four principles that emerged from that work, and how you might apply them to your own workflows. &lt;/span&gt;&lt;/p&gt;
&lt;h3&gt;&lt;strong style="vertical-align: baseline;"&gt;Principle 1: Verify delegated work &lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;If we are to permit AI to delegate tasks, we want it to do more than arbitrarily assign work. Agents should intelligently break down work into tasks that can be reliably verified. In our research, we call this "contract-first decomposition." &lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Like human delegation, this takes thoughtful deliberation. With people, this might mean a leader understanding their team’s strengths, and perhaps checking their work before it’s completed. For AI, there’s a similar learning curve. The orchestrating AI (the manager that sits atop a multi-agentic system) may consider multiple plans for how best to decompose and assign work, and keep decomposing sub-goals into smaller and smaller chunks until they become sufficiently simple to monitor and verify. Ideally, this should result in a plan where everything can be reliably graded. In reality, however, this may not always be possible to achieve. &lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Sometimes, it may be necessary to involve subjective assessment of whether work has been completed successfully, in line with expectations. Rather than being a problem, identifying such components helps us determine where human time is best spent, and how best to involve human expert judgement in oversight of agentic systems.&lt;/span&gt;&lt;/p&gt;
&lt;h3&gt;&lt;strong style="vertical-align: baseline;"&gt;Principle 2: Be smart about cost&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;The research framework helps us answer a question that keeps coming up with customers: Can this particular task be handled by a smaller, cheaper model? Enterprises are increasingly attentive to cost, and rightly so. &lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Finding the right balance between performance and budget is tricky. Taking a complex problem, like payroll, and handing it off to a lightweight model, might not be powerful enough for the results you want. On the other hand, it’s unnecessary to route simple tasks, like reformatting a spreadsheet, to a strong reasoning model. &lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;According to the research, an agent that is intelligent about delegation would learn to recognize these scenarios, and match each task to the right tool or endpoint, to achieve the desired result and maximum reliability at a minimum cost. Use of model routing capabilities within API gateways is becoming a popular choice among customers, in addition to the alternative for using client-side proxies (such as LiteLLM). &lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;You can learn more about model routing &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/api-gateway/docs/model-routing-overview"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;here&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;.&lt;/span&gt;&lt;/p&gt;
&lt;h3&gt;&lt;strong style="vertical-align: baseline;"&gt;Principle #3: Respect sensitive data &lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Many workflows handle private, sensitive data, and AI agents need to respect those boundaries and permissions. For example, if you’re deploying your orchestrator agent for payroll data, you know that agent should never pass along its full set of information to a sub-agent. This not only compromises security, but also bloats the context window for agents and degrades performance. An agent should grant the absolute minimum permissions required to complete that specific assignment, and nothing more. &lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;The challenging part arises when needing to demonstrate, according to our first principle, that work has been reliably completed, without revealing private information. According to the research , advanced cryptography can help address this, via techniques such as zero-knowledge proofs. Zero-knowledge proofs enable one AI agent to prove to the other AI agent that a planned computation was performed correctly, without revealing the data itself. For example, an agent tasked with analyzing a sensitive dataset can generate a succinct non-interactive argument of knowledge that proves a specific property of the result. This enables the delegator to instantly verify the validity of the proof.&lt;/span&gt;&lt;/p&gt;
&lt;h3&gt;&lt;strong style="vertical-align: baseline;"&gt;Principles #4: Beware the zone of indifference&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;The zone of indifference is a term coined by Chester Barnard, an American business executive, in his 1938 book called &lt;/span&gt;&lt;span style="font-style: italic; vertical-align: baseline;"&gt;The Function of the Executive. &lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;The zone is the space in which an employee will accept a task without questioning it. The task usually falls within their scope, so they unconsciously accept it. For example, if you’re a sales rep and your manager asks you to attend an upcoming pitch with a valued client, you probably wouldn’t push back or think too deeply about it.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;As expressed in the research, current AI systems are defined by post-training safety filters and system instructions. As long as a request does not trigger a hard violation, the model complies. But when considering the emerging agentic web, this compliance might actually create a systemic risk. As mentioned in the research, “As delegation chains lengthen (? → ? → ?), a broad zone of indifference allows subtle intent mismatches or context-dependent harms to propagate rapidly downstream, with each agent acting as an unthinking router rather than a responsible actor.”&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;This has serious implications, because it means intelligent delegation requires “dynamic cognitive friction.” This means validating the information provided to agents to ensure that they are accurate, relevant, controlled and efficient. &lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;This way, an agent can recognize when a request is ambiguous enough to warrant stepping outside their zone of indifference to challenge the delegator, or request human verification. Human participation and oversight similarly presume a degree of cognitive friction and active engagement, though this must be carefully managed, so as not to over-burden the users of the system. Human time is valuable and should only be invoked when necessary.&lt;/span&gt;&lt;/p&gt;
&lt;h3&gt;&lt;strong style="vertical-align: baseline;"&gt;Looking ahead&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;At Google Cloud, our long-term goal is to integrate agents naturally and efficiently into organizations, which will mean delegating to and from human experts and respecting boundaries. Together, we believe this will deliver business value beyond what individual agents can handle. &lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Ready to navigate the agentic web? Read Google DeepMind’s report paper, &lt;/span&gt;&lt;a href="https://arxiv.org/abs/2602.11865" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Intelligent AI Delegation&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;, on arXiv. &lt;/span&gt;&lt;/p&gt;
&lt;hr/&gt;
&lt;p&gt;&lt;sub&gt;&lt;span style="font-style: italic; vertical-align: baseline;"&gt;Note: A special thanks to Matija Franklin, Simon Osindero from Google DeepMind, and Vishal Agarwal, Andrea Morange from Google Cloud, for their contributions.&lt;/span&gt;&lt;/sub&gt;&lt;/p&gt;&lt;/div&gt;</description><pubDate>Fri, 21 Aug 2026 16:00:00 +0000</pubDate><guid>https://cloud.google.com/blog/products/ai-machine-learning/how-agents-can-delegate-better/</guid><category>AI &amp; Machine Learning</category><og xmlns:og="http://ogp.me/ns#"><type>article</type><title>How agents can delegate better</title><description></description><site_name>Google</site_name><url>https://cloud.google.com/blog/products/ai-machine-learning/how-agents-can-delegate-better/</url></og><author xmlns:author="http://www.w3.org/2005/Atom"><name>Nenad Tomasev</name><title>Research Scientist, Google DeepMind</title><department></department><company></company></author><author xmlns:author="http://www.w3.org/2005/Atom"><name>Reshu Yadav</name><title>Applied AI Blackbelt, Google Cloud</title><department></department><company></company></author></item><item><title>Expanding Google Antigravity for enterprise customers</title><link>https://cloud.google.com/blog/products/ai-machine-learning/expanding-google-antigravity-for-enterprise-customers/</link><description>&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Since announcing Google Antigravity in Gemini Enterprise Agent Platform at I/O in May, we’ve heard helpful feedback from our customers. Your developers want easy access to coding agents across surfaces. Your enterprise governance team wants security controls and license management. And your finance team wants pooled usage so that no prepaid token ever goes unused. Now, everybody finally gets what they want:&lt;/span&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;span style="vertical-align: baseline;"&gt;Antigravity is available now as part of &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/gemini/enterprise/docs/ai-developer-tools-overview"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;eligible&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; Gemini Enterprise app subscriptions, including out-of-the-box administrative and spend controls.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;span style="vertical-align: baseline;"&gt;New &lt;/span&gt;&lt;a href="https://antigravity.google/blog/antigravity-ide-extensions" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;IDE extensions&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; let developers use Antigravity in the IDEs of their choice, including &lt;a href="https://marketplace.visualstudio.com/items?itemName=Google.google-antigravity" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;VS Code&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;.&lt;/span&gt;&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;Unify AI developer tools and enterprise-grade controls in one subscription&lt;/span&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Equipping your developers with advanced agentic tools shouldn't mean managing separate add-on licenses, invoices, billing consoles or security settings. With AI developer tools included i&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;n Gemini Enterprise subscriptions, administrators can easily &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/gemini/enterprise/docs/ai-developer-tools-overview"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;enable&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; Antigravity and &lt;/span&gt;&lt;a href="http://d.android.com/gemini-in-android" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Android Studio&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; for users with &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/gemini/enterprise/docs/ai-developer-tools-overview"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;eligible&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; Gemini Enterprise Standard, Plus, and Standard Emerging Market licenses, and maintain full governance with spend, security, observability and usage metrics consolidated in the Gemini Enterprise admin console.&lt;/span&gt;&lt;/p&gt;
&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;Unblock your developers while controlling spend&lt;/span&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;With billing flexibility and cost management tools in Gemini Enterprise, you can ensure your developers have the resources they need while managing costs:&lt;/span&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;a href="https://docs.cloud.google.com/billing/docs/how-to/budgets-spend-caps"&gt;&lt;strong style="text-decoration: underline; vertical-align: baseline;"&gt;Granular spend thresholds&lt;/strong&gt;&lt;/a&gt;&lt;strong style="vertical-align: baseline;"&gt;:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Administrators can set monthly project-level budget caps directly in the Billing console, with additional per-user and team controls rolling out later this year.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;a href="https://docs.cloud.google.com/gemini/enterprise/docs/quotas-and-overages"&gt;&lt;strong style="text-decoration: underline; vertical-align: baseline;"&gt;Pooled quotas&lt;/strong&gt;&lt;/a&gt;&lt;strong style="vertical-align: baseline;"&gt;: &lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;&lt;span style="vertical-align: baseline;"&gt;Shared token pools provide flexibility to high-demand teams, preventing purchased quota from sitting idle across the organization. &lt;/span&gt;&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;a href="https://docs.cloud.google.com/gemini/enterprise/docs/configure-overages"&gt;&lt;strong style="text-decoration: underline; vertical-align: baseline;"&gt;Overage enablement&lt;/strong&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;: To maintain continuous developer workflows when pooled quotas are met, administrators can opt into overages with monthly spend caps, smoothly transitioning excess usage to standard consumption-based rates. &lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;a href="https://docs.cloud.google.com/gemini/enterprise/docs/ai-developer-tools-metrics"&gt;&lt;strong style="text-decoration: underline; vertical-align: baseline;"&gt;Usage metrics&lt;/strong&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;: Centralized usage tracking provides visibility into token consumption, API calls, and developer activity, enabling organizations to continuously optimize their AI investments.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;&lt;/div&gt;
&lt;div class="block-image_full_width"&gt;






  
    &lt;div class="article-module h-c-page"&gt;
      &lt;div class="h-c-grid"&gt;
  

    &lt;figure class="article-image--large
      
      
        h-c-grid__col
        h-c-grid__col--6 h-c-grid__col--offset-3
        
        
      "
      &gt;

      
      
        
        &lt;img
            src="https://storage.googleapis.com/gweb-cloudblog-publish/original_images/Gif_1_vK8C2Nf.gif"
        
          alt="Gif 1"&gt;
        
        &lt;/a&gt;
      
    &lt;/figure&gt;

  
      &lt;/div&gt;
    &lt;/div&gt;
  




&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;Safeguard your organization’s code and data with built-in privacy and security&lt;/span&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Gemini Enterprise subscriptions bring Google Antigravity under Google Cloud’s standard security and compliance protections. Administrators and IT teams can set clear boundaries around workspace access, enable full audit logging, and enforce data privacy from a single console:&lt;/span&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;a href="https://docs.cloud.google.com/gemini/enterprise/docs/ai-developer-tools-settings"&gt;&lt;strong style="text-decoration: underline; vertical-align: baseline;"&gt;Configurable security policies&lt;/strong&gt;&lt;/a&gt;&lt;strong style="vertical-align: baseline;"&gt;:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; &lt;span style="vertical-align: baseline;"&gt;Enforce security and compliance controls, such as workspace sandboxing, and browser and MCP server access, to help ensure AI agents operate safely within authorized enterprise environments.&lt;/span&gt;&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;a href="https://docs.cloud.google.com/gemini/enterprise/docs/ai-developer-tools-settings"&gt;&lt;strong style="text-decoration: underline; vertical-align: baseline;"&gt;Central audit logging&lt;/strong&gt;&lt;/a&gt;&lt;strong style="vertical-align: baseline;"&gt;:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Enable comprehensive audit logging with a single toggle, capturing prompts, agent responses, and metadata for compliance reporting.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;a href="https://cloud.google.com/privacy?e=48754805"&gt;&lt;strong style="text-decoration: underline; vertical-align: baseline;"&gt;Data privacy&lt;/strong&gt;&lt;/a&gt;&lt;strong style="vertical-align: baseline;"&gt;: &lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;Maintain data ownership under Google Cloud’s Terms of Service, ensuring all agent activity executes strictly within your secure cloud boundary.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;&lt;/div&gt;
&lt;div class="block-video"&gt;



&lt;div class="article-module article-video "&gt;
  &lt;figure&gt;
    &lt;a class="h-c-video h-c-video--marquee"
      href="https://youtube.com/watch?v=CJnchBVryCQ"
      data-glue-modal-trigger="uni-modal-CJnchBVryCQ-"
      data-glue-modal-disabled-on-mobile="true"&gt;

      
        

        &lt;div class="article-video__aspect-image"
          style="background-image: url(https://storage.googleapis.com/gweb-cloudblog-publish/images/maxresdefault_R43tOCG.max-1000x1000.jpg);"&gt;
          &lt;span class="h-u-visually-hidden"&gt;Antigravity in Gemini Enterprise - AI Developer Tools Settings&lt;/span&gt;
        &lt;/div&gt;
      
      &lt;svg role="img" class="h-c-video__play h-c-icon h-c-icon--color-white"&gt;
        &lt;use xlink:href="#mi-youtube-icon"&gt;&lt;/use&gt;
      &lt;/svg&gt;
    &lt;/a&gt;

    
  &lt;/figure&gt;
&lt;/div&gt;

&lt;div class="h-c-modal--video"
     data-glue-modal="uni-modal-CJnchBVryCQ-"
     data-glue-modal-close-label="Close Dialog"&gt;
   &lt;a class="glue-yt-video"
      data-glue-yt-video-autoplay="true"
      data-glue-yt-video-height="99%"
      data-glue-yt-video-vid="CJnchBVryCQ"
      data-glue-yt-video-width="100%"
      href="https://youtube.com/watch?v=CJnchBVryCQ"
      ng-cloak&gt;
   &lt;/a&gt;
&lt;/div&gt;

&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;Bring agentic coding directly into your team’s preferred development environments &lt;/span&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Starting today your developers can use Antigravity across the surfaces they already know and use — including &lt;/span&gt;&lt;a href="https://marketplace.visualstudio.com/items?itemName=Google.google-antigravity" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Visual Studio Code&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;, &lt;/span&gt;&lt;a href="https://marketplace.visualstudio.com/items?itemName=Google.GoogleAntigravity" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Visual Studio&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; (preview), Jetbrains (preview) and Zed IDEs (preview)  via the &lt;/span&gt;&lt;a href="https://antigravity.google/blog/antigravity-ide-extensions" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;new IDE extensions&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; as well as the Antigravity 2.0 desktop app, and the Antigravity CLI. &lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Throughout all surfaces, administrators can enforce corporate identity standards while removing setup friction for technical teams via native support for Workforce Identity Federation (WIF) and Application Default Credentials (ADC).&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-image_full_width"&gt;






  
    &lt;div class="article-module h-c-page"&gt;
      &lt;div class="h-c-grid"&gt;
  

    &lt;figure class="article-image--large
      
      
        h-c-grid__col
        h-c-grid__col--6 h-c-grid__col--offset-3
        
        
      "
      &gt;

      
      
        
        &lt;img
            src="https://storage.googleapis.com/gweb-cloudblog-publish/original_images/VSCode_IDE_Plugin_gif_v3_720p.gif"
        
          alt="VSCode IDE Plugin"&gt;
        
        &lt;/a&gt;
      
        &lt;figcaption class="article-image__caption "&gt;&lt;p data-block-key="6topl"&gt;AGY IDE extension demo&lt;/p&gt;&lt;/figcaption&gt;
      
    &lt;/figure&gt;

  
      &lt;/div&gt;
    &lt;/div&gt;
  




&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;What our customers are saying&lt;/span&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;From rapid code generation to end-to-end task automation, Google Antigravity is giving engineering teams the momentum of cutting edge AI development backed by the stability, governance, and scale of Google Cloud. Here is how leading enterprise customers and partners are driving measurable outcomes in production:&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;“Deploying Antigravity in Gemini Enterprise allows Accenture to arm our engineers with Google DeepMind’s premier technology on the secure, trusted foundation of Google Cloud. Abstracting away operational complexity ensures our teams don't have to choose between developer speed and enterprise-grade governance — freeing them to deliver high-velocity engineering and transformative value for our clients.” — Chetna Sehgal, Global Practice Lead, &lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;Accenture Google Business Group&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;“At AirAsia and across the Group, we’re all about empowering our people. Bringing highly capable Gemini models directly into our daily workflows with Antigravity 2.0 does exactly that. We are putting the most advanced AI capabilities into the hands of our entire workforce, from software engineering to finance, marketing, legal, HR and much more. This empowers both our developers and critical back-office teams to innovate at an unprecedented pace and drive proven time-savings across the board.” — Nikunj Shanti, CTO, &lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;AirAsia Next&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;“Enterprises are moving beyond AI experimentation and expecting measurable business outcomes. With Gemini Enterprise and next-generation developer tools like Antigravity 2.0 and Antigravity CLI, we see significant opportunities to further embed agentic AI, particularly the advanced reasoning capabilities of Gemini models, directly into software delivery workflows. This goes beyond productivity as it enables faster decision-making, higher code quality, and reduced technical debt at scale. What stands out is how these capabilities are helping our teams evolve from writing code to orchestrating outcomes, strengthening every phase of the software development lifecycle while scaling innovation securely and responsibly.” — Rakesh Aerath, President, Asia Pacific Global Delivery Centers of Excellence, &lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;CGI&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;“Every developer workflow is unique, and agentic AI should adapt to the engineer, not the other way around. With Google Antigravity supported across developers' preferred IDEs, the desktop app, and the CLI, Cognizant can seamlessly embed agentic engineering across our global delivery centers. It gives our teams the freedom to choose their preferred surface while delivering high-velocity, secure software for our clients.”— Rajesh Varrier, President, Global Operations and Chairman &amp;amp; Managing Director, &lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;Cognizant India&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;“Antigravity has played a key role in advancing our AI-first strategy at Datamatics. Over the last few months, our teams have used it to rapidly build and deploy multiple applications, accelerate solution development, and embed AI into core business processes. From AI Impact Hub to analytics and sales enablement solutions, it has helped us move beyond experimentation to real execution, delivering measurable business outcomes while enabling teams to innovate faster and at scale.”— Vijay Venkatachalam, Vice President, Information Systems Group (ISG), &lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;Datamatics&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;"Embedding Google Antigravity's autonomous capabilities directly into Gemini Enterprise development environments allows our internal and Forward Deployed Engineering teams to automate complex tasks. Supported by the platform’s new FinOps and governance controls, our teams can confidently focus on orchestrating high-value, secure and cost-efficient outcomes for our clients at scale."  — Faruk Muratovic, US AI &amp;amp; Engineering Strategy and Services Leader,&lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt; Deloitte&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;“Adopting Antigravity places Wipro at the leading edge of the AI-driven software development lifecycle. Combining Antigravity 2.0, CLI, and the new IDE extensions with a seamless developer experience and superior code accuracy fits naturally into our AI-first engineering strategy. Working with Google Cloud allows us to accelerate software delivery and bring next-generation value to our global enterprise clients.”  — Debashish Ghosh, Vice President and Global Head, Google Partnership, &lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;Wipro&lt;/strong&gt;&lt;/p&gt;
&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;Get started with Antigravity in Gemini Enterprise&lt;/span&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Google Antigravity in Gemini Enterprise is available today for &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/gemini/enterprise/docs/ai-developer-tools-overview"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;eligible&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; Gemini Enterprise Standard, Plus, and Standard Emerging Market licenses, with broader support coming soon.&lt;/span&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;For administrators:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Visit the &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/gemini/enterprise/docs/ai-developer-tools-overview"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;enterprise setup guide&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; to enable AI Developer tools for Google Antigravity and Android Studio. &lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;For developers: &lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;Start &lt;/span&gt;&lt;a href="https://antigravity.google/docs/home" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;building&lt;/span&gt;&lt;/a&gt; &lt;span style="vertical-align: baseline;"&gt;with Antigravity 2.0, the Antigravity CLI, or your preferred IDEs via Antigravity IDE extensions.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;&lt;/div&gt;</description><pubDate>Thu, 20 Aug 2026 17:30:00 +0000</pubDate><guid>https://cloud.google.com/blog/products/ai-machine-learning/expanding-google-antigravity-for-enterprise-customers/</guid><category>AI &amp; Machine Learning</category><media:content height="540" url="https://storage.googleapis.com/gweb-cloudblog-publish/images/antigravity_enterprise.max-600x600.jpg" width="540"></media:content><og xmlns:og="http://ogp.me/ns#"><type>article</type><title>Expanding Google Antigravity for enterprise customers</title><description></description><image>https://storage.googleapis.com/gweb-cloudblog-publish/images/antigravity_enterprise.max-600x600.jpg</image><site_name>Google</site_name><url>https://cloud.google.com/blog/products/ai-machine-learning/expanding-google-antigravity-for-enterprise-customers/</url></og><author xmlns:author="http://www.w3.org/2005/Atom"><name>Scott Densmore</name><title>Senior Director of Engineering</title><department></department><company></company></author><author xmlns:author="http://www.w3.org/2005/Atom"><name>Damith Karunaratne</name><title>Group Product Manager</title><department></department><company></company></author></item><item><title>How AlloyDB ScaNN scales vector search to 10 billion vectors</title><link>https://cloud.google.com/blog/products/databases/alloydb-scann-index-four-level-tree-improves-vector-search/</link><description>&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;To satisfy the demands of enterprise-grade agentic AI applications, underlying vector databases often struggle to scale effectively as modern use cases can scale to billions of vectors.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;As a fully managed PostgreSQL-compatible database service, &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/alloydb/docs/overview"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;AlloyDB&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; is engineered to handle demanding enterprise workloads. Combining Google's infrastructure with the reliability of commercial databases, it delivers high availability, scalability, and includes a cutting-edge analytical engine, optimal for agentic AI use cases. A key part of this is its &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/alloydb/docs/ai/create-scann-index"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;ScaNN index&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;, which now operates efficiently &lt;/span&gt;&lt;span style="font-style: italic; vertical-align: baseline;"&gt;at a scale of 10 billion vectors&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;. This was achieved through a major architectural enhancement: an innovative &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/alloydb/docs/ai/create-scann-index#four-level-tree-index"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;four-level tree (preview)&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; paired with efficient memory usage&lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;.&lt;/strong&gt;&lt;/p&gt;
&lt;h3&gt;&lt;strong style="vertical-align: baseline;"&gt;The 10 billion vector scale challenge&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Scaling to a 10 billion vector workload presents significant memory and computational challenges. Previous AlloyDB ScaNN tree-based index was limited to &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/alloydb/docs/ai/create-scann-index#two-level-tree-index"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;two&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;- or &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/alloydb/docs/ai/create-scann-index#three-level-tree-index"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;three&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;-level tree configurations, and attempting to scale those structures led to several bottlenecks:&lt;/span&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Increased compute intensity:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Larger tree structures demand significantly more operations for both index construction and query traversal.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Memory constraints:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; The sampling processes required for 10 billion vectors can easily exceed the system's available memory capacity.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h3&gt;&lt;strong style="vertical-align: baseline;"&gt;Solution: Four-level architecture&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;The introduction of a &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/alloydb/docs/ai/create-scann-index#four-level-tree-index"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;four-level tree (preview)&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; is the primary innovation in the recent AlloyDB ScaNN release. This architecture, illustrated in Figure 1, employs a top-down strategy to optimize the balance between accuracy and build efficiency. To maintain high performance and mitigate recall loss, the system integrates key enhancements such as Top-K branch, &lt;/span&gt;&lt;a href="https://research.google/blog/soar-new-algorithms-for-even-faster-vector-search-with-scann/#:~:text=ScaNN%20is%20open%2Dsourced%20on%20GitHub%20and%20can%20be%20easily%20installed%20via%20Pip." rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;SOAR&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;, &lt;/span&gt;&lt;a href="https://arxiv.org/abs/1908.10396" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;centroid adjustment&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; and balanced tree shape.&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-image_full_width"&gt;






  
    &lt;div class="article-module h-c-page"&gt;
      &lt;div class="h-c-grid"&gt;
  

    &lt;figure class="article-image--large
      
      
        h-c-grid__col
        h-c-grid__col--6 h-c-grid__col--offset-3
        
        
      "
      &gt;

      
      
        
        &lt;img
            src="https://storage.googleapis.com/gweb-cloudblog-publish/images/1_fpfICUj.max-1000x1000.jpg"
        
          alt="1"&gt;
        
        &lt;/a&gt;
      
        &lt;figcaption class="article-image__caption "&gt;&lt;p data-block-key="uveoq"&gt;Figure 1. AlloyDB ScaNN four-level tree architecture&lt;/p&gt;&lt;/figcaption&gt;
      
    &lt;/figure&gt;

  
      &lt;/div&gt;
    &lt;/div&gt;
  




&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;This design has two primary benefits:&lt;/span&gt;&lt;/p&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;1. Reduced compute intensity via hierarchical partitioning&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;The four-level architecture drastically reduces compute intensity by using hierarchical partitioning to restrict the volume of vectors scanned during a query. Instead of traversing a flat or poorly segmented space, the multi-layered hierarchy narrows down the search path exponentially. Figure 2 illustrates the search spaces across different tree levels, demonstrating how structural layering optimizes traversal efficiency:&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-image_full_width"&gt;






  
    &lt;div class="article-module h-c-page"&gt;
      &lt;div class="h-c-grid"&gt;
  

    &lt;figure class="article-image--large
      
      
        h-c-grid__col
        h-c-grid__col--6 h-c-grid__col--offset-3
        
        
      "
      &gt;

      
      
        
        &lt;img
            src="https://storage.googleapis.com/gweb-cloudblog-publish/images/2_LWwXC70.max-1000x1000.jpg"
        
          alt="2"&gt;
        
        &lt;/a&gt;
      
        &lt;figcaption class="article-image__caption "&gt;&lt;p data-block-key="uveoq"&gt;Figure 2. Search space for two-, three- and four-level trees&lt;/p&gt;&lt;/figcaption&gt;
      
    &lt;/figure&gt;

  
      &lt;/div&gt;
    &lt;/div&gt;
  




&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Two-level:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Utilizes coarse partitioning to guide queries, resulting in a basic search complexity of &lt;/span&gt;&lt;em&gt;&lt;span style="vertical-align: baseline;"&gt;O(&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;N&lt;/span&gt;&lt;sup&gt;&lt;span style="vertical-align: baseline;"&gt;1/2&lt;/span&gt;&lt;/sup&gt;&lt;/em&gt;&lt;span style="vertical-align: baseline;"&gt;&lt;em&gt;)&lt;/em&gt;&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Three-level:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Introduces an intermediate layer to further subdivide clusters, narrowing exploration to &lt;/span&gt;&lt;em&gt;&lt;span style="vertical-align: baseline;"&gt;O(&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;N&lt;/span&gt;&lt;sup&gt;&lt;span style="vertical-align: baseline;"&gt;1/3&lt;/span&gt;&lt;/sup&gt;&lt;/em&gt;&lt;span style="vertical-align: baseline;"&gt;&lt;em&gt;)&lt;/em&gt;&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Four-level:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Implements refined, highly granular partitions that optimize traversal efficiency down to &lt;/span&gt;&lt;em&gt;&lt;span style="vertical-align: baseline;"&gt;O(&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;N&lt;/span&gt;&lt;sup&gt;&lt;span style="vertical-align: baseline;"&gt;1/4&lt;/span&gt;&lt;/sup&gt;&lt;/em&gt;&lt;span style="vertical-align: baseline;"&gt;&lt;em&gt;)&lt;/em&gt;&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;, sufficiently allowing for more than 10-billion vectors.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;By dynamically expanding hierarchical layers as the dataset expands, AlloyDB ScaNN maintains ultra-low query latency and avoids computational scale walls from impacting performance.&lt;/span&gt;&lt;/p&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;2. Efficient memory usage&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Achieving a 10 billion vector scale requires high memory efficiency. AlloyDB ScaNN uses these strategies to maximize memory management performance:&lt;/span&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Balanced tree shape construction:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; The four-level tree utilizes a &lt;/span&gt;&lt;span style="font-style: italic; vertical-align: baseline;"&gt;balanced&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt; configuration to circumvent memory limitations that restrict the size of training datasets. This balanced architecture effectively leverages reduced sampling sizes to construct high-fidelity tree partitions.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Sampling optimization:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; When the system encounters memory limitations, it generates a condensed sampling set that considers performance and accuracy. &lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h3&gt;&lt;strong style="vertical-align: baseline;"&gt;Performance test results&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;By leveraging the innovative four-level tree architecture in our internal tests, we are able to achieve the following performance results:&lt;/span&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;span style="vertical-align: baseline;"&gt;AlloyDB can scale to over 10 billion vectors with its ScaNN index.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;span style="vertical-align: baseline;"&gt;AlloyDB can deliver &amp;lt;= 51 ms p95 latency and 95% recall at 10 billion vectors with its ScaNN index.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h3&gt;&lt;strong style="vertical-align: baseline;"&gt;Get started today&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Experience AlloyDB ScaNN's &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/alloydb/docs/ai/create-scann-index#four-level-tree-index"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;four-level tree (preview)&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; architecture today. You can deploy ScaNN for AlloyDB by following our &lt;/span&gt;&lt;a href="https://cloud.google.com/alloydb/docs/quickstart/create-and-connect"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;quickstart guide&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; to set up an instance. For optimized, high-speed vector search, refer to the &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/alloydb/docs/ai/create-scann-index"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;official ScaNN documentation&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;. New users can also &lt;/span&gt;&lt;a href="https://console.cloud.google.com/alloydb/create-trial-cluster?_gl=1*qsd2cd*_up*MQ..&amp;amp;gclid=CjwKCAjwooq3BhB3EiwAYqYoEh91xxGzv4xrmyMJJ_BPfF4X8cv-I3kINwvnMI2pADozFQPsrHnaOhoCbioQAvD_BwE&amp;amp;gclsrc=aw.ds"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;explore AlloyDB&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; through our &lt;/span&gt;&lt;a href="https://cloud.google.com/blog/products/databases/run-your-postgresql-database-in-an-alloydb-free-trial-cluster"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;30-day free trial&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; program. We can’t wait to hear about what you build!&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;</description><pubDate>Thu, 20 Aug 2026 16:00:00 +0000</pubDate><guid>https://cloud.google.com/blog/products/databases/alloydb-scann-index-four-level-tree-improves-vector-search/</guid><category>AI &amp; Machine Learning</category><category>Databases</category><og xmlns:og="http://ogp.me/ns#"><type>article</type><title>How AlloyDB ScaNN scales vector search to 10 billion vectors</title><description></description><site_name>Google</site_name><url>https://cloud.google.com/blog/products/databases/alloydb-scann-index-four-level-tree-improves-vector-search/</url></og><author xmlns:author="http://www.w3.org/2005/Atom"><name>Bin Song</name><title>Software Engineer</title><department></department><company></company></author><author xmlns:author="http://www.w3.org/2005/Atom"><name>Itai Rosenblatt</name><title>Engineering Manager</title><department></department><company></company></author></item><item><title>10 questions every startup should answer before moving to production with their AI prototype</title><link>https://cloud.google.com/blog/topics/developers-practitioners/10-questions-for-your-startup-developers/</link><description>&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;It’s never been easier to start an AI-powered startup on Google Cloud. &lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;You grab an API key from &lt;/span&gt;&lt;a href="https://aistudio.google.com/" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Google AI Studio&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; at breakfast, paste it into Antigravity, and by lunch you’ll have a nascent prototype of your product.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;But it’s not all one straight line to progress. It's common to bump into these three challenges as you build out your stack:&lt;/span&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;span style="vertical-align: baseline;"&gt;A leaked API key racks up a&lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt; &lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;large bill in 48 hours&lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;.&lt;/strong&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;span style="vertical-align: baseline;"&gt;A "quick" migration from AI Studio to &lt;/span&gt;&lt;a href="https://cloud.google.com/vertex-ai"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Gemini Enterprise Agent Platform&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; stalls the roadmap for weeks because nobody on the team owns Identity and Access Management (IAM).&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;span style="vertical-align: baseline;"&gt;The launch works, until the app starts returning &lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;HTTP 429 Too Many Requests&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt; because of default per-project quotas, and there's no clean path to more capacity without paying a premium.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;None of these are unique edge cases. . They're  default failure modes of moving fast without a plan, and we've all done it at least once.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Below are the 10 questions every startup should be ready to answer &lt;/span&gt;&lt;span style="font-style: italic; vertical-align: baseline;"&gt;before&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt; they scale,  grouped into the three phases where decisions can shape your future: &lt;/span&gt;&lt;/p&gt;
&lt;ol&gt;
&lt;li aria-level="1" style="list-style-type: decimal; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Onboard&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; (setting up your own projects and identities right)&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: decimal; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Scale&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; (getting more throughput without breaking the bank) &lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: decimal; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Govern&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; (keeping costs, keys, and agents from running away). &lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;These ten are scoped to the prototype-to-production transition itself. Each question ends with a short, runnable snippet you can copy into your own project today. Adjacent decisions that matter just as much but aren't specific to that move, your data layer and RAG architecture, CI/CD, network design, are deliberately out of frame here. &lt;/span&gt;&lt;/p&gt;
&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;Onboard: get the foundation right (in the first hour).&lt;/span&gt;&lt;/h3&gt;
&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;#1 Where should I start: Google AI Studio or Gemini Enterprise Agent Platform?&lt;/span&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Both surfaces expose the same Gemini family of models, but they solve different problems.&lt;/span&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Google AI Studio&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; (with the Gemini Developer API) is the fastest path from an idea to working code. A browser IDE, an API key, a generous free tier, and no cloud project to configure. It's where most ideas should start, and Google's own guidance says as much.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Gemini Enterprise Agent Platform &lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;(formerly Vertex AI) has the same Gemini models (plus 3rd party and OSS ones)  with enterprise controls around them: IAM and service-account auth instead of raw keys, VPC Service Controls, Cloud Logging and Monitoring, reserved capacity, regional endpoints, and the compliance surface your first enterprise customer's security review will ask about.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;The right answer for most startups is &lt;/span&gt;&lt;span style="font-style: italic; vertical-align: baseline;"&gt;both&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;, &lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;sequenced deliberately&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;: &lt;/span&gt;&lt;span style="font-style: italic; vertical-align: baseline;"&gt;first&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt; prototype in AI Studio, &lt;/span&gt;&lt;span style="font-style: italic; vertical-align: baseline;"&gt;then&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt; migrate &lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;before&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; you have real users. The danger for startups is treating them as interchangeable solutions, AI Studio's simple key model does not translate to enterprise controls, and Agent Platform's IAM model might look like overkill until the day it saves you from a stolen-credential incident.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;It's less work than it sounds like.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;The unified &lt;/span&gt;&lt;a href="https://github.com/googleapis/python-genai" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;google-genai&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; SDK targets both:&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-code"&gt;&lt;dl&gt;
    &lt;dt&gt;code_block&lt;/dt&gt;
    &lt;dd&gt;&amp;lt;ListValue: [StructValue([(&amp;#x27;code&amp;#x27;, &amp;#x27;# Prototype: Google AI Studio, raw API key\r\nfrom google import genai\r\nclient = genai.Client(api_key=&amp;quot;YOUR_AI_STUDIO_KEY&amp;quot;)\r\n\r\n# Production: GEAP, no key — uses Application Default Credentials (ADC)\r\nfrom google import genai\r\nclient = genai.Client(\r\n    vertexai=True,\r\n    project=&amp;quot;my-startup-prod&amp;quot;,\r\n    location=&amp;quot;us-central1&amp;quot;,\r\n)\r\n\r\nresp = client.models.generate_content(\r\n    model=&amp;quot;gemini-2.5-pro&amp;quot;,\r\n    contents=&amp;quot;Summarize this contract in three bullets.&amp;quot;,\r\n)\r\nprint(resp.text)&amp;#x27;), (&amp;#x27;language&amp;#x27;, &amp;#x27;&amp;#x27;), (&amp;#x27;caption&amp;#x27;, &amp;lt;wagtail.rich_text.RichText object at 0x7f1876cfc880&amp;gt;)])]&amp;gt;&lt;/dd&gt;
&lt;/dl&gt;&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;#2  How do I set up a Google Cloud project without becoming an IAM expert?&lt;/span&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;The biggest reason startups stall on the migration to Agent Platform isn't the code, it's the operational leap from "here's an API key" to a cloud project with folders, service accounts, org policies, logging, and IAM bindings. If your team doesn't have a dedicated cloud admin, that first project setup can eat a week of engineering time. &lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Three moves cut that dramatically:&lt;/span&gt;&lt;/p&gt;
&lt;ol&gt;
&lt;li aria-level="1" style="list-style-type: decimal; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Use an opinionated project template instead of clicking through the console.&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; The &lt;/span&gt;&lt;a href="https://console.cloud.google.com/cloud-setup"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Cloud Setup checklist&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; and the &lt;/span&gt;&lt;a href="https://cloud.google.com/architecture/framework"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Google Cloud Architecture Framework&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; give you a production-grade folder hierarchy (prod / non-prod / dev), a central logging + monitoring project, &lt;/span&gt;&lt;a href="https://cloud.google.com/security-command-center"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Security Command Center&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; turned on, and baseline org policies, without you having to design them from scratch.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: decimal; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Enable the APIs you'll actually use, once.&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Batch it so you're not doing it project-by-project when you need it. The billing-link step is not optional. Every paid API you're about to enable will refuse to activate on a project with no billing account attached, so we handle that first.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: decimal; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;span style="vertical-align: baseline;"&gt;Let &lt;/span&gt;&lt;a href="https://cloud.google.com/iam/docs/role-picker-gemini"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Gemini pick the roles&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;, but ask it for the narrow ones. You don't have to memorize the roles reference. In the Grant access dialog, Help me choose roles lets you describe the task in plain language, "this service account needs to call Gemini models and read one Cloud Storage bucket", and get predefined roles back with the reasoning shown. &lt;/span&gt;&lt;span style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, Oxygen, Ubuntu, Cantarell, 'Open Sans', 'Helvetica Neue', sans-serif;"&gt;One catch worth knowing on day one: by default it suggests roles that cover common journeys, which usually means a service's Admin, Editor, or Viewer. Those are broader than you want. Say "least privileged" or "narrowest access" in the prompt and it returns granular roles instead. Same amount of typing, considerably smaller blast radius when a credential leaks.&lt;br/&gt;&lt;br/&gt;&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;Sources: &lt;/span&gt;&lt;a href="https://cloud.google.com/iam/docs/role-picker-gemini" style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, Oxygen, Ubuntu, Cantarell, 'Open Sans', 'Helvetica Neue', sans-serif;"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Get predefined role suggestions with Gemini assistance&lt;/span&gt;&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ol&gt;&lt;/div&gt;
&lt;div class="block-code"&gt;&lt;dl&gt;
    &lt;dt&gt;code_block&lt;/dt&gt;
    &lt;dd&gt;&amp;lt;ListValue: [StructValue([(&amp;#x27;code&amp;#x27;, &amp;#x27;# One-shot: create a Vertex-ready project and turn on the services a\r\n# typical AI startup uses.\r\ngcloud projects create my-startup-prod --name=&amp;quot;My Startup (prod)&amp;quot;\r\ngcloud config set project my-startup-prod\r\n\r\n# REQUIRED before enabling billing-dependent APIs (aiplatform, run, etc.).\r\n# Use `gcloud billing accounts list` to find your billing account ID.\r\ngcloud billing projects link my-startup-prod --billing-account=012345-6789AB-CDEF01\r\n\r\ngcloud services enable \\\r\n  aiplatform.googleapis.com \\\r\n  run.googleapis.com \\\r\n  artifactregistry.googleapis.com \\\r\n  logging.googleapis.com \\\r\n  monitoring.googleapis.com \\\r\n  secretmanager.googleapis.com \\\r\n  cloudbilling.googleapis.com&amp;#x27;), (&amp;#x27;language&amp;#x27;, &amp;#x27;&amp;#x27;), (&amp;#x27;caption&amp;#x27;, &amp;lt;wagtail.rich_text.RichText object at 0x7f1876cfc160&amp;gt;)])]&amp;gt;&lt;/dd&gt;
&lt;/dl&gt;&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Sources: &lt;/span&gt;&lt;a href="https://cloud.google.com/sdk/gcloud/reference/services/enable"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;gcloud services enable reference&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;, · &lt;/span&gt;&lt;a href="https://cloud.google.com/sdk/gcloud/reference/billing/projects/link"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;gcloud billing projects link (GA)&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;,  &lt;/span&gt;&lt;a href="https://cloud.google.com/vertex-ai/docs/start/cloud-environment"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;GE Agent Platform environment setup&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;If you're a solo founder, resist the urge to build in your personal GCP account. Create a proper organization or self-owned org first, then create the project &lt;/span&gt;&lt;span style="font-style: italic; vertical-align: baseline;"&gt;inside&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt; it. That single decision can make everything else, fromIAM to billing and audit, dramatically easier.&lt;/span&gt;&lt;/p&gt;
&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;#3 I'm on Google Cloud, how should my code actually authenticate: API keys, service accounts, or user credentials?&lt;/span&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;There's a hierarchy of safety here, and the easiest option is rarely the right one in production.&lt;/span&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Raw API keys&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; are fine for local prototyping. They are dangerous in production because they are long-lived, easy to leak into a client bundle or a public repo, and grant unbounded access until you notice.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;User credentials via OAuth (application default credentials)&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; are best for interactive tools, CLIs, and any code that runs on a developer's laptop.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Service accounts with least-privilege IAM roles&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; are the right answer for anything running on a server, in a container, or in a scheduled job.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;The pattern you're aiming for is one where your &lt;/span&gt;&lt;span style="font-style: italic; vertical-align: baseline;"&gt;code never sees a key at all&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;. It just calls the &lt;/span&gt;&lt;a href="https://google-auth.readthedocs.io/en/latest/" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Google Auth library&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;, which quietly reads Application Default Credentials (ADC) from the environment,  a short-lived token minted for whichever service account is attached to your Cloud Run service, GKE workload, or Compute Engine VM. You get enterprise-grade auth without writing any auth code.&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-code"&gt;&lt;dl&gt;
    &lt;dt&gt;code_block&lt;/dt&gt;
    &lt;dd&gt;&amp;lt;ListValue: [StructValue([(&amp;#x27;code&amp;#x27;, &amp;#x27;# On a developer laptop\r\ngcloud auth application-default login\r\n\r\n# On a server (Cloud Run, GKE, etc.) — no login, no key file.\r\n# Attach a service account with just the roles the app needs.\r\ngcloud run deploy my-agent \\\r\n  --image=us-docker.pkg.dev/my-startup-prod/agents/api:v1 \\\r\n  --service-account=agent-runtime@my-startup-prod.iam.gserviceaccount.com \\\r\n  --region=us-central1&amp;#x27;), (&amp;#x27;language&amp;#x27;, &amp;#x27;&amp;#x27;), (&amp;#x27;caption&amp;#x27;, &amp;lt;wagtail.rich_text.RichText object at 0x7f1876cfc1f0&amp;gt;)])]&amp;gt;&lt;/dd&gt;
&lt;/dl&gt;&lt;/div&gt;
&lt;div class="block-code"&gt;&lt;dl&gt;
    &lt;dt&gt;code_block&lt;/dt&gt;
    &lt;dd&gt;&amp;lt;ListValue: [StructValue([(&amp;#x27;code&amp;#x27;, &amp;#x27;# Application code — notice: no keys, no secrets.\r\nfrom google import genai\r\n\r\nclient = genai.Client(\r\n    vertexai=True,\r\n    project=&amp;quot;my-startup-prod&amp;quot;,\r\n    location=&amp;quot;us-central1&amp;quot;,\r\n)&amp;#x27;), (&amp;#x27;language&amp;#x27;, &amp;#x27;&amp;#x27;), (&amp;#x27;caption&amp;#x27;, &amp;lt;wagtail.rich_text.RichText object at 0x7f1876e15430&amp;gt;)])]&amp;gt;&lt;/dd&gt;
&lt;/dl&gt;&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Do one last favor to your future self: give that service account the &lt;/span&gt;&lt;span style="font-style: italic; vertical-align: baseline;"&gt;minimum&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt; IAM role your workload actually needs,  usually &lt;/span&gt;&lt;a href="https://cloud.google.com/vertex-ai/docs/general/access-control"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;roles/aiplatform.user&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; for calling models, not the broader admin roles. It takes an extra 30 seconds and prevents the credential from becoming a master key if it leaks.&lt;/span&gt;&lt;/p&gt;
&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;#4 When should I actually stop procrastinating and migrate from AI Studio's API key to Agent Platform's IAM model?&lt;/span&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Sooner than you'd like,  and the correct trigger is &lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;not&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; when it breaks. It's when any of these is true:&lt;/span&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;span style="vertical-align: baseline;"&gt;Your key has left your laptop (checked into a repo, pasted into a Slack, shipped in a mobile app).&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;span style="vertical-align: baseline;"&gt;You have more than one person on the team who needs to call the API.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;span style="vertical-align: baseline;"&gt;You're spending more than a few hundred dollars a month.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;span style="vertical-align: baseline;"&gt;You're about to onboard paying customers.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;A potential pitfall that can catch growing startups off guard is simple: a leaked Gemini API key on an account that normally spends $180 a month gets scraped from a public repo and used to run distillation attacks,  accumulating tens of thousands of dollars in charges before the owner even sees the first billing alert. The &lt;/span&gt;&lt;a href="https://cloud.google.com/architecture/framework/security/shared-responsibility-shared-fate"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Google Cloud Shared Responsibility Model&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; is unambiguous: the customer is liable for charges incurred with their own valid credentials.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;The migration itself is genuinely smaller than the anxiety around it. In &lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;google-genai&lt;/strong&gt;&lt;strong style="vertical-align: baseline;"&gt; &lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;it's the two-line change shown in #1. What takes real time is the &lt;/span&gt;&lt;span style="font-style: italic; vertical-align: baseline;"&gt;project setup&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt; around it, which is exactly why #2 exists.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Practical checklist for cutover day:&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-code"&gt;&lt;dl&gt;
    &lt;dt&gt;code_block&lt;/dt&gt;
    &lt;dd&gt;&amp;lt;ListValue: [StructValue([(&amp;#x27;code&amp;#x27;, &amp;#x27;# 1. Revoke every existing AI Studio key that has ever left a laptop.\r\n#    (Go to https://aistudio.google.com/apikey and delete them.)\r\n\r\n# 2. Confirm your production code has no api_key= arguments.\r\ngrep -rn &amp;quot;api_key&amp;quot; src/\r\n\r\n# 3. Enable GEAP and confirm ADC works locally.\r\ngcloud services enable aiplatform.googleapis.com\r\ngcloud auth application-default login\r\npython -c &amp;quot;\r\nfrom google import genai\r\nc = genai.Client(vertexai=True, project=\&amp;#x27;my-startup-prod\&amp;#x27;, location=\&amp;#x27;us-central1\&amp;#x27;)\r\nprint(c.models.generate_content(model=\&amp;#x27;gemini-2.5-flash\&amp;#x27;, contents=\&amp;#x27;ping\&amp;#x27;).text)\r\n&amp;quot;&amp;#x27;), (&amp;#x27;language&amp;#x27;, &amp;#x27;&amp;#x27;), (&amp;#x27;caption&amp;#x27;, &amp;lt;wagtail.rich_text.RichText object at 0x7f1876e15160&amp;gt;)])]&amp;gt;&lt;/dd&gt;
&lt;/dl&gt;&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;If step 3 prints a response, you're on Agent Platform.&lt;/span&gt;&lt;/p&gt;
&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;&lt;span style="vertical-align: baseline;"&gt;Scale: get more capacity without paying a premium.&lt;/span&gt;&lt;/span&gt;&lt;/h3&gt;
&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;#5 Now that I'm shipping, why on earth am I getting all these &lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;HTTP 429&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt; errors, and how do I make them stop?&lt;/span&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;429 Too Many Requests&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt; from Agent Platform almost always means one of two things:&lt;/span&gt;&lt;/p&gt;
&lt;ol&gt;
&lt;li aria-level="1" style="list-style-type: decimal; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;span style="vertical-align: baseline;"&gt;You've hit the &lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;Dynamic Shared Quota (DSQ)&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; ceiling for your project's tier. DSQ is a shared pool sized against your project's history,  new projects start with modest limits by design, to prevent abuse across the platform.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: decimal; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;span style="vertical-align: baseline;"&gt;You're calling a &lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;global endpoint&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; during a global demand spike, competing with worldwide traffic for shared capacity.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;The instinctive reaction is to file a quota-increase ticket. You can do that if you must,  but two architectural moves usually solve the problem faster and cheaper.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;Pin to a regional endpoint.&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Over half of startup traffic on Agent Platform defaults to global routing. Pinning to a specific region (say &lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;us-central1&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;) sidesteps global contention and typically improves latency at the same time. (One narrow exception, which we'll get to in the next question: if you specifically want Priority PayGo, that feature currently only ships on the `global` endpoint. For everything else, pin regionally.):&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-code"&gt;&lt;dl&gt;
    &lt;dt&gt;code_block&lt;/dt&gt;
    &lt;dd&gt;&amp;lt;ListValue: [StructValue([(&amp;#x27;code&amp;#x27;, &amp;#x27;from google import genai\r\n\r\n# Global (default): competes against worldwide demand.\r\n# Regional: routes only to the regional cluster, less contention.\r\nclient = genai.Client(\r\n    vertexai=True,\r\n    project=&amp;quot;my-startup-prod&amp;quot;,\r\n    location=&amp;quot;us-central1&amp;quot;,   # &amp;lt;-- this is the one-line fix\r\n)&amp;#x27;), (&amp;#x27;language&amp;#x27;, &amp;#x27;&amp;#x27;), (&amp;#x27;caption&amp;#x27;, &amp;lt;wagtail.rich_text.RichText object at 0x7f1876d1ba60&amp;gt;)])]&amp;gt;&lt;/dd&gt;
&lt;/dl&gt;&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;Add real retry and backoff.&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; A 429 is a retryable signal, not a fatal error. Any production client should have exponential backoff with jitter. The modern google-genai SDK &lt;/span&gt;&lt;a href="https://cloud.google.com/vertex-ai/docs/reference/rest#retry_settings"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;ships this behavior&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; built in, but only if you actually enable it. This is easy to overlook. Don't reach for the classic `google.api_core.retry.if_transient_error` decorator you may have seen on older Vertex code. It's designed for the legacy exception classes and does not recognize the new `google.genai.errors.APIError,  so it will silently pass 429s through without retrying. Use the SDK's built-in retry options instead:&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-code"&gt;&lt;dl&gt;
    &lt;dt&gt;code_block&lt;/dt&gt;
    &lt;dd&gt;&amp;lt;ListValue: [StructValue([(&amp;#x27;code&amp;#x27;, &amp;#x27;from google import genai\r\nfrom google.genai import types\r\n\r\nclient = genai.Client(\r\n    vertexai=True, project=&amp;quot;my-startup-prod&amp;quot;, location=&amp;quot;us-central1&amp;quot;,\r\n    http_options=types.HttpOptions(retry_options=types.HttpRetryOptions(\r\n        attempts=5, initial_delay=1.0, max_delay=60.0, exp_base=2.0, jitter=1.0,\r\n        http_status_codes=[408, 429, 500, 502, 503, 504],\r\n    ))\r\n)&amp;#x27;), (&amp;#x27;language&amp;#x27;, &amp;#x27;&amp;#x27;), (&amp;#x27;caption&amp;#x27;, &amp;lt;wagtail.rich_text.RichText object at 0x7f1876e28c10&amp;gt;)])]&amp;gt;&lt;/dd&gt;
&lt;/dl&gt;&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;How do you see this coming?  &lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;Preferably not from a user telling you. Agent Platform publishes serving metrics to Cloud Monitoring, and there is a prebuilt dashboard you don't have to assemble: Console → Agent Platform → Dashboard → Model observability. It gives you requests per second, token throughput, first-token latency, and error rates out of the box.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;The metric to actually alert on is &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;aiplatform.googleapis.com/publisher/online_serving/model_invocation_count&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt;. It carries an &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;error_category&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; label with values of &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;user&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt;, &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;system&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt;, or &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;capacity&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt;. Alerting on &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;capacity&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; isolates genuine throttling from your own bad requests, which a raw 429 count won't do.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;One thing worth internalizing, because it trips people up: you cannot build a "warn me at 80% of my quota" alert for Standard PayGo. Under Dynamic Shared Quota there is no fixed per-project number to be at 80% of. A 429 means transient contention for shared capacity, not that you crossed a line. Percent-of-limit alerting only becomes meaningful once you're on Provisioned Throughput, which does expose real limit metrics.&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-code"&gt;&lt;dl&gt;
    &lt;dt&gt;code_block&lt;/dt&gt;
    &lt;dd&gt;&amp;lt;ListValue: [StructValue([(&amp;#x27;code&amp;#x27;, &amp;#x27;gcloud monitoring policies create --policy-from-file=capacity-alert.yaml&amp;#x27;), (&amp;#x27;language&amp;#x27;, &amp;#x27;&amp;#x27;), (&amp;#x27;caption&amp;#x27;, &amp;lt;wagtail.rich_text.RichText object at 0x7f18841977c0&amp;gt;)])]&amp;gt;&lt;/dd&gt;
&lt;/dl&gt;&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Sources: &lt;/span&gt;&lt;a href="https://cloud.google.com/monitoring/api/metrics_gcp_a_b"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Agent Platform metrics list&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;, &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/gemini-enterprise-agent-platform/models/model-observability"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Model observability dashboard&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;, &lt;/span&gt;&lt;a href="https://github.com/googleapis/python-genai/blob/main/google/genai/types.py" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;RetryOptions source&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;,  &lt;/span&gt;&lt;a href="https://github.com/googleapis/google-cloud-python/blob/main/packages/google-api-core/google/api_core/retry/retry_base.py" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;core retry_base.py&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;, &lt;/span&gt;&lt;a href="https://github.com/googleapis/python-genai/blob/main/google/genai/errors.py" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;genai &lt;/span&gt;&lt;/a&gt;&lt;a href="http://errors.py" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;errors.py&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;,  &lt;/span&gt;&lt;a href="https://cloud.google.com/blog/products/ai-machine-learning/reduce-429-errors-on-vertex-ai?e=48754805"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;reduce 429 errors&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;, &lt;/span&gt;&lt;a href="https://cloud.google.com/sdk/gcloud/reference/monitoring/policies/create"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;gcloud monitoring policies create&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;, &lt;/span&gt;&lt;a href="https://cloud.google.com/vertex-ai/generative-ai/docs/dynamic-shared-quota"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Dynamic Shared Quota&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Follow the &lt;/span&gt;&lt;a href="https://cloud.google.com/vertex-ai/generative-ai/docs/quotas"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Agent Platform rate limits documentation&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; to understand what&lt;/span&gt;&lt;span style="font-style: italic; vertical-align: baseline;"&gt; your&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt; project's current ceiling actually is before you assume you've outgrown it. &lt;/span&gt;&lt;/p&gt;
&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;#6 Which consumption mode do I pay for: Standard PayGo, Priority PayGo, or Provisioned Throughput? &lt;/span&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Three consumption models, three completely different workload shapes, and three completely different ways to proceed. Picking the right one can help startups see meaningful savings on AI bills. First let’s define them and then see when they are, or aren’t, a good fit: &lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;Standard PayGo (DSQ)&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;: Pay per token from a shared pool; cheap, no guarantees.&lt;br/&gt;&lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;Priority PayGo&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;: Pay per token at a premium to jump the queue.&lt;br/&gt;&lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;Provisioned Throughput (PT)&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;: Prepay for reserved capacity; predictable, use it or lose it.&lt;br/&gt;&lt;br/&gt;&lt;/span&gt;&lt;/p&gt;
&lt;div align="left"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;&lt;table&gt;&lt;colgroup&gt;&lt;col/&gt;&lt;col/&gt;&lt;col/&gt;&lt;/colgroup&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th scope="col" style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;Consumption type&lt;/strong&gt;&lt;/p&gt;
&lt;/th&gt;
&lt;th scope="col" style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;Best for&lt;/strong&gt;&lt;/p&gt;
&lt;/th&gt;
&lt;th scope="col" style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;Watch out for&lt;/strong&gt;&lt;/p&gt;
&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;Standard PayGo (DSQ)&lt;/strong&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Early-stage, low-QPS, spiky prototype traffic&lt;/span&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;429s during spikes; no reliability SLO&lt;/span&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;a href="https://cloud.google.com/vertex-ai/generative-ai/docs/priority-paygo"&gt;&lt;strong style="text-decoration: underline; vertical-align: baseline;"&gt;Priority PayGo&lt;/strong&gt;&lt;/a&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Bursty, revenue-critical traffic that can't tolerate 429s&lt;/span&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Roughly 1.8x the standard token price&lt;/span&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;a href="https://cloud.google.com/vertex-ai/generative-ai/docs/provisioned-throughput"&gt;&lt;strong style="text-decoration: underline; vertical-align: baseline;"&gt;Provisioned Throughput (PT)&lt;/strong&gt;&lt;/a&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Steady, predictable, high-volume production traffic&lt;/span&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Wasted spend if utilization is under ~40%; overflow to PayGo on spikes&lt;/span&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;The dominant startup mistake is buying PT too early. Usually  this happens the  week after a big launch when it feels like traffic will only ever go up. PT is &lt;/span&gt;&lt;span style="font-style: italic; vertical-align: baseline;"&gt;reserved&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt; capacity. You  pay whether you use it or not, and it only starts paying you back once your baseline is genuinely predictable, not just aspirational.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Here’s a pragmatic sequence:&lt;/span&gt;&lt;/p&gt;
&lt;ol&gt;
&lt;li aria-level="1" style="list-style-type: decimal; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Weeks one through four on Standard PayGo.&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Use it to measure your real request shape (tokens per minute at p50 and p99, request bursts, batchable vs. real-time split).&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: decimal; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;When you get your first bad 429 storm,&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; flip on Priority PayGo for the traffic that actually matters. It's a config change, not a purchase order,  nobody in procurement needs to be involved:&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ol&gt;&lt;/div&gt;
&lt;div class="block-code"&gt;&lt;dl&gt;
    &lt;dt&gt;code_block&lt;/dt&gt;
    &lt;dd&gt;&amp;lt;ListValue: [StructValue([(&amp;#x27;code&amp;#x27;, &amp;#x27;# Priority PayGo request: use the global endpoint + two extra headers.\r\nfrom google import genai\r\nfrom google.genai import types\r\n\r\nclient = genai.Client(vertexai=True, project=&amp;quot;my-startup-prod&amp;quot;, location=&amp;quot;global&amp;quot;)\r\nresp = client.models.generate_content(\r\n    model=&amp;quot;gemini-2.5-pro&amp;quot;,\r\n    contents=&amp;quot;Rank these support tickets by urgency: ...&amp;quot;,\r\n    config=types.GenerateContentConfig(\r\n        # Priority PayGo headers, per current GEAP docs.\r\n        http_options=types.HttpOptions(headers={&amp;quot;X-Vertex-AI-LLM-Request-Type&amp;quot;: &amp;quot;shared&amp;quot;, &amp;quot;X-Vertex-AI-LLM-Shared-Request-Type&amp;quot;: &amp;quot;priority&amp;quot;}),\r\n    ),\r\n)&amp;#x27;), (&amp;#x27;language&amp;#x27;, &amp;#x27;&amp;#x27;), (&amp;#x27;caption&amp;#x27;, &amp;lt;wagtail.rich_text.RichText object at 0x7f1884c311f0&amp;gt;)])]&amp;gt;&lt;/dd&gt;
&lt;/dl&gt;&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;p role="presentation"&gt;&lt;span style="vertical-align: baseline;"&gt;3. &lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;Once you can predict your baseline TPM,&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; buy PT to cover the flat baseline and let anything above it overflow to PayGo. That's the &lt;/span&gt;&lt;a href="https://cloud.google.com/blog/products/ai-machine-learning/provisioned-throughput-on-vertex-ai"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;combined pattern Google recommends&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; for exactly this reason. Best of both worlds, not marketing spin.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt; Sources: &lt;/span&gt;&lt;a href="https://cloud.google.com/vertex-ai/generative-ai/docs/priority-paygo"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Priority PayGo docs&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;, &lt;/span&gt;&lt;a href="https://github.com/googleapis/python-genai/blob/main/google/genai/types.py" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;google-genai HttpOptions source&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;, &lt;/span&gt;&lt;a href="https://cloud.google.com/vertex-ai/generative-ai/docs/reference/rest"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;GEAP REST reference&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;. &lt;/span&gt;&lt;/p&gt;
&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;#7 Which of my requests actually need to be live, and which should be batch jobs?&lt;/span&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Most startup workloads are secretly batch jobs pretending to be real-time. Every one you move off the interactive path frees up DSQ headroom for the traffic that genuinely needs to be fast,  the traffic where a user is actually watching a spinner.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Three questions to help you sort your traffic:&lt;/span&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;span style="vertical-align: baseline;"&gt;Does a human have to see the result within a second? That means:  &lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;Live inference.&lt;/strong&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;span style="vertical-align: baseline;"&gt;Can the user wait a few seconds and see a spinner? That means:  Still live, but a candidate for streaming.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;span style="vertical-align: baseline;"&gt;Would the user tolerate "we'll email you when it's ready" or "check back in a bit"?  That means: &lt;/span&gt;&lt;a href="https://cloud.google.com/vertex-ai/generative-ai/docs/multimodal/batch-prediction"&gt;&lt;strong style="text-decoration: underline; vertical-align: baseline;"&gt;Batch prediction&lt;/strong&gt;&lt;/a&gt;&lt;strong style="vertical-align: baseline;"&gt;.&lt;/strong&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Batch prediction on Agent Platform runs in a completely separate queue, does not consume your interactive DSQ, and is typically about half the price of on-demand inference. That's a rare double win: faster live traffic &lt;/span&gt;&lt;span style="font-style: italic; vertical-align: baseline;"&gt;and&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt; a lower bill.&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-code"&gt;&lt;dl&gt;
    &lt;dt&gt;code_block&lt;/dt&gt;
    &lt;dd&gt;&amp;lt;ListValue: [StructValue([(&amp;#x27;code&amp;#x27;, &amp;#x27;# Kick off a batch prediction job from a JSONL file in Cloud Storage.\r\n# Each line is one prompt; results land in another Cloud Storage prefix.\r\nfrom google import genai\r\nfrom google.genai import types\r\n\r\nclient = genai.Client(vertexai=True, project=&amp;quot;my-startup-prod&amp;quot;, location=&amp;quot;us-central1&amp;quot;)\r\n\r\njob = client.batches.create(\r\n    model=&amp;quot;gemini-2.5-flash&amp;quot;,\r\n    src=&amp;quot;gs://my-startup-prod-batch/inputs/nightly-summaries.jsonl&amp;quot;,\r\n    config=types.CreateBatchJobConfig(\r\n        dest=&amp;quot;gs://my-startup-prod-batch/outputs/&amp;quot;,\r\n    ),\r\n)\r\nprint(job.name, job.state)&amp;#x27;), (&amp;#x27;language&amp;#x27;, &amp;#x27;&amp;#x27;), (&amp;#x27;caption&amp;#x27;, &amp;lt;wagtail.rich_text.RichText object at 0x7f1884c312b0&amp;gt;)])]&amp;gt;&lt;/dd&gt;
&lt;/dl&gt;&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Common candidates: nightly document summarization, background classification of new signups, bulk translation, embedding backfills, evaluation runs against your test set. If any of those are on your live path today, moving them is often the single highest-leverage change you can make this week.&lt;/span&gt;&lt;/p&gt;
&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;Govern: Keep costs, keys, and agents under control.&lt;/span&gt;&lt;/h3&gt;
&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;#8 How do I set spend caps that actually reduce cost, and not just send me polite emails while my bill triples?&lt;/span&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Until recently the honest answer was that budgets only notify, and you had to build your own brake pedal. That changed in July. There are now three mechanisms, and you should think of them as layers.&lt;/span&gt;&lt;/p&gt;
&lt;ol&gt;
&lt;li aria-level="1" style="list-style-type: decimal; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;span style="vertical-align: baseline;"&gt;A &lt;/span&gt;&lt;a href="https://cloud.google.com/billing/docs/how-to/budgets-spend-caps"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;spend cap budget&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; (Preview). Cloud Billing budgets can now enforce rather than just email. Set a spend cap on a project and, when usage costs cross 100% of the budget, Google pauses the service until you manually lift it. Agent Platform is explicitly on the eligible list, alongside the Gemini API, Cloud Run, and Cloud Run functions. Alerts still fire at 50% and 80%, so the pause isn't a surprise.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Three things to know before you rely on it:&lt;/span&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;span style="vertical-align: baseline;"&gt;Each cap covers one project and one eligible service. It is not account-wide protection. If you want Agent Platform and Cloud Run both capped, that's two caps. &lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;span style="vertical-align: baseline;"&gt;Enforcement is not instant and is based on estimated costs. Overages past the cap are billed as normal, so set the number below your real ceiling. Lifting it is manual, and service resumption can take up to an hour. It also pauses Provisioned Throughput usage, so if you've prepaid for capacity, a cap hit stops that too.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;span style="vertical-align: baseline;"&gt;It's in Preview as of publication, and the eligible-service list is documented as growing. Check the current list before you design around it.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p role="presentation"&gt;&lt;span style="vertical-align: baseline;"&gt;2. &lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;A billing budget with a Pub/Sub trigger that disables billing&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;. Still the right tool when you need blast radius the spend cap can't give you: multiple services at once, an entire project, or a service that isn't eligible yet. When the budget hits a threshold, Pub/Sub fires a Cloud Function that detaches the billing account, which stops all billable activity within minutes. Blunter and more dangerous than the native cap — it can leave resources unrecoverable — so reach for it second, not first. &lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;Full walkthrough: &lt;/span&gt;&lt;a href="https://cloud.google.com/billing/docs/how-to/notify"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Automatically respond to budget notifications&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;.&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-code"&gt;&lt;dl&gt;
    &lt;dt&gt;code_block&lt;/dt&gt;
    &lt;dd&gt;&amp;lt;ListValue: [StructValue([(&amp;#x27;code&amp;#x27;, &amp;#x27;# Sketch: create a budget SCOPED TO ONE PROJECT that publishes to Pub/Sub at 50%, 90%, 100%.\r\ngcloud billing budgets create \\\r\n  --billing-account=012345-6789AB-CDEF01 \\\r\n  --display-name=&amp;quot;my-startup-prod hard stop&amp;quot; \\\r\n  --budget-amount=2000USD \\\r\n  --filter-projects=projects/my-startup-prod \\\r\n  --threshold-rule=percent=0.5 \\\r\n  --threshold-rule=percent=0.9 \\\r\n  --threshold-rule=percent=1.0,basis=current-spend \\\r\n  --notifications-rule-pubsub-topic=projects/my-startup-prod/topics/budget-alerts&amp;#x27;), (&amp;#x27;language&amp;#x27;, &amp;#x27;&amp;#x27;), (&amp;#x27;caption&amp;#x27;, &amp;lt;wagtail.rich_text.RichText object at 0x7f18771a8e50&amp;gt;)])]&amp;gt;&lt;/dd&gt;
&lt;/dl&gt;&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Sources: &lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt; &lt;/span&gt;&lt;a href="https://cloud.google.com/billing/docs/how-to/budgets-spend-caps"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Manage spend cap budgets&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;, &lt;/span&gt;&lt;a href="https://cloud.google.com/billing/docs/how-to/budgets-programmatic-notifications"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Set up programmatic notifications&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; &lt;/span&gt;&lt;a href="https://cloud.google.com/sdk/gcloud/reference/billing/budgets/create"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;gcloud billing budgets create reference&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;, &lt;/span&gt;&lt;a href="https://cloud.google.com/billing/docs/how-to/budgets"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Cloud Billing budgets concepts&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;, &lt;/span&gt;&lt;a href="https://cloud.google.com/billing/docs/how-to/disable-billing-with-notifications"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Disable billing with notifications walkthrough&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;, &lt;/span&gt;&lt;a href="https://cloud.google.com/billing/docs/how-to/budgets-programmatic-notifications"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Programmatic notification payload schema&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Two things to get ahead of  for, as the defaults can cause unexpected issues: &lt;/span&gt;&lt;/p&gt;
&lt;ol&gt;
&lt;li aria-level="1" style="list-style-type: decimal; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;span style="vertical-align: baseline;"&gt;Limit your budget scope: Without --filter-projects, your budget applies to your entire billing account. A spike in any project will trigger the kill switch for everything. &lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: decimal; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;span style="vertical-align: baseline;"&gt;Deploy locally: The budget notification doesn't specify which project is affected. To ensure the kill switch only affects the intended project, deploy your Cloud Function in the same project you're protecting (e.g., my-startup-prod).&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Then wire up a tiny Cloud Function to that topic that calls &lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;projects.updateBillingInfo&lt;/strong&gt;&lt;strong style="vertical-align: baseline;"&gt; &lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;to unlink the billing account when the 100% threshold fires. That is your circuit breaker.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;Mechanical ceilings via quota overrides.&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Even if you never set up the above kill switch, you can cap the &lt;/span&gt;&lt;span style="font-style: italic; vertical-align: baseline;"&gt;rate&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt; at which cost can accumulate by setting explicit per-model, per-region quotas below the platform default. If your app never legitimately needs more than 500 requests per minute for &lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;gemini-2.5-pro&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;, cap it there in the &lt;/span&gt;&lt;a href="https://cloud.google.com/docs/quotas/view-manage"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Cloud Quotas console&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;; a leaked key can't burn what the quota flatly refuses to serve.&lt;/span&gt;&lt;/p&gt;
&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;#9 Where should I actually keep secrets? (Not in .env files!)&lt;/span&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;The short answer is: &lt;/span&gt;&lt;a href="https://cloud.google.com/secret-manager"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Secret Manager&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;. Not  in environment variables, not in &lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;.env&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; files, and never in your repo. Grant read access via IAM only to the service account that needs it.&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-code"&gt;&lt;dl&gt;
    &lt;dt&gt;code_block&lt;/dt&gt;
    &lt;dd&gt;&amp;lt;ListValue: [StructValue([(&amp;#x27;code&amp;#x27;, &amp;#x27;# Store a third-party API key (Stripe, OpenAI, whatever).\r\necho -n &amp;quot;sk_live_xxx&amp;quot; | gcloud secrets create stripe-live-key --data-file=-\r\n\r\n# Grant only the runtime service account access to read it.\r\ngcloud secrets add-iam-policy-binding stripe-live-key \\\r\n  --member=serviceAccount:agent-runtime@my-startup-prod.iam.gserviceaccount.com \\\r\n  --role=roles/secretmanager.secretAccessor&amp;#x27;), (&amp;#x27;language&amp;#x27;, &amp;#x27;&amp;#x27;), (&amp;#x27;caption&amp;#x27;, &amp;lt;wagtail.rich_text.RichText object at 0x7f18771a8130&amp;gt;)])]&amp;gt;&lt;/dd&gt;
&lt;/dl&gt;&lt;/div&gt;
&lt;div class="block-code"&gt;&lt;dl&gt;
    &lt;dt&gt;code_block&lt;/dt&gt;
    &lt;dd&gt;&amp;lt;ListValue: [StructValue([(&amp;#x27;code&amp;#x27;, &amp;#x27;# Application code fetches it at startup; nothing lives on disk.\r\nfrom google.cloud import secretmanager\r\nsm = secretmanager.SecretManagerServiceClient()\r\nresp = sm.access_secret_version(\r\n    name=&amp;quot;projects/my-startup-prod/secrets/stripe-live-key/versions/latest&amp;quot;\r\n)\r\nstripe_key = resp.payload.data.decode(&amp;quot;utf-8&amp;quot;)&amp;#x27;), (&amp;#x27;language&amp;#x27;, &amp;#x27;&amp;#x27;), (&amp;#x27;caption&amp;#x27;, &amp;lt;wagtail.rich_text.RichText object at 0x7f1876ec7fa0&amp;gt;)])]&amp;gt;&lt;/dd&gt;
&lt;/dl&gt;&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Then two little disciplines that pay for themselves the first time you need them:&lt;/span&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Rotation on a schedule and on suspicion.&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Secret Manager versions are cheap; treat them as immutable and roll forward. &lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Detection when a secret leaks.&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; &lt;/span&gt;&lt;a href="https://cloud.google.com/secret-manager/docs/event-notifications"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Secret Manager notifications&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; and Google Cloud's &lt;/span&gt;&lt;a href="https://cloud.google.com/sensitive-data-protection"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Sensitive Data Protection&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; can catch keys checked into a repo or pasted into a log stream,  before an attacker does.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;For any AI application that acts on a user's behalf, calls Gmail on their behalf, reads a Drive folder, hits a third-party SaaS with the user's credentials, do &lt;/span&gt;&lt;span style="font-style: italic; vertical-align: baseline;"&gt;not&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt; store a long-lived token. Use OAuth 2.0 with short-lived access tokens and a refresh flow, so that when a user rage-quits or a compromised account gets revoked, the agent loses access at the same time. &lt;/span&gt;&lt;/p&gt;
&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;#10  How do I stop my brand new AI agent from doing something it absolutely shouldn't?&lt;/span&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;An agent that can call tools, browse the web, or execute code needs the same defense-in-depth thinking as any other production service, arguably more, because it makes decisions that neither you nor the model can fully predict in advance.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Four layers, none optional once you have real users:&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;1. Identity for the agent itself.&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Give the agent its own service account, scoped only to the resources and tools it genuinely needs,  the exact same least-privilege principle as any other workload. Agent Engine supports first-class &lt;/span&gt;&lt;a href="https://cloud.google.com/vertex-ai/generative-ai/docs/agent-engine/identity"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;agent identity&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; so every action can be attributed to a specific agent instance in your audit logs.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;2. Sandboxed code execution.&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; If your agent runs generated code,  a common pattern for data-analysis or "run this Python for me" flows, do not run it in your application process. Use an &lt;/span&gt;&lt;a href="https://cloud.google.com/vertex-ai/generative-ai/docs/code-execution"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;isolated sandbox&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; so a bad combination can't touch your production data.&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-code"&gt;&lt;dl&gt;
    &lt;dt&gt;code_block&lt;/dt&gt;
    &lt;dd&gt;&amp;lt;ListValue: [StructValue([(&amp;#x27;code&amp;#x27;, &amp;#x27;# Enable server-side code execution inside a sandbox for a request.\r\nfrom google import genai\r\nfrom google.genai import types\r\n\r\nclient = genai.Client(vertexai=True, project=&amp;quot;my-startup-prod&amp;quot;, location=&amp;quot;us-central1&amp;quot;)\r\nresp = client.models.generate_content(\r\n    model=&amp;quot;gemini-2.5-pro&amp;quot;,\r\n    contents=&amp;quot;Compute the correlation between these two columns: ...&amp;quot;,\r\n    config=types.GenerateContentConfig(\r\n        tools=[types.Tool(code_execution=types.ToolCodeExecution())],\r\n    ),\r\n)&amp;#x27;), (&amp;#x27;language&amp;#x27;, &amp;#x27;&amp;#x27;), (&amp;#x27;caption&amp;#x27;, &amp;lt;wagtail.rich_text.RichText object at 0x7f1876ec7af0&amp;gt;)])]&amp;gt;&lt;/dd&gt;
&lt;/dl&gt;&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;3. Prompt and response filtering.&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; &lt;/span&gt;&lt;a href="https://cloud.google.com/security-command-center/docs/model-armor-overview"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Model Armor&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; sits in front of your model calls and screens for prompt injection, jailbreaks, sensitive-data exfiltration, and off-brand output,  all of which are essentially guaranteed the moment you have real users being real users.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;4. Behavioral monitoring.&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; &lt;/span&gt;&lt;a href="https://cloud.google.com/security-command-center"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Security Command Center&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; with &lt;/span&gt;&lt;a href="https://cloud.google.com/security-command-center/docs/concepts-security-sources"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;threat detection&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; flags anomalies in agent behavior,  a service account suddenly calling an API it's never touched before, an agent reaching out to an unfamiliar external host, an unexpected spike in privileged operations. In near-real-time.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;None of these are optional once your agent is acting on behalf of a real user or handling real money.&lt;/span&gt;&lt;/p&gt;
&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;Your homework, so to speak:&lt;/span&gt;&lt;/h3&gt;
&lt;ol&gt;
&lt;li aria-level="1" style="list-style-type: decimal; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;span style="vertical-align: baseline;"&gt;Audit for raw API keys in your repo, your notebooks, and your production runtime. Rotate anything that shouldn't be there.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: decimal; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;span style="vertical-align: baseline;"&gt;Move any workload that doesn't need a synchronous response to the Batch API.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: decimal; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;span style="vertical-align: baseline;"&gt;Turn on the Model observability dashboard and put one alert on capacity errors, so the next 429 reaches you before it reaches a customer.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: decimal; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;span style="vertical-align: baseline;"&gt;Set a spend cap on the project, and keep an eye out for 50% and 80% alerts. If usage crosses 100% of the budget, Google will pause the service until you manually lift it.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Do those four things this week and you're already ahead of most startups shipping AI features. &lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="font-style: italic; vertical-align: baseline;"&gt;Have a scenario you'd like us to cover next? Reach us at &lt;/span&gt;&lt;a href="https://cloud.google.com/startup"&gt;&lt;span style="font-style: italic; text-decoration: underline; vertical-align: baseline;"&gt;Google Cloud for Startups&lt;/span&gt;&lt;/a&gt;&lt;span style="font-style: italic; vertical-align: baseline;"&gt;.&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;</description><pubDate>Thu, 20 Aug 2026 16:00:00 +0000</pubDate><guid>https://cloud.google.com/blog/topics/developers-practitioners/10-questions-for-your-startup-developers/</guid><category>AI &amp; Machine Learning</category><category>Startups</category><category>Developers &amp; Practitioners</category><og xmlns:og="http://ogp.me/ns#"><type>article</type><title>10 questions every startup should answer before moving to production with their AI prototype</title><description></description><site_name>Google</site_name><url>https://cloud.google.com/blog/topics/developers-practitioners/10-questions-for-your-startup-developers/</url></og><author xmlns:author="http://www.w3.org/2005/Atom"><name>Sergio Villani</name><title>Technical Solutions, Google Cloud AI</title><department></department><company></company></author></item><item><title>How Box is unlocking multimodal enterprise agents with Gemini Embeddings 2</title><link>https://cloud.google.com/blog/topics/partners/box-ai-agents-gemini-embeddings-multimodal-enterprise-ai/</link><description>&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Enterprise content management is experiencing its biggest architectural shift since the cloud migration era. &lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;For years, enterprises have stored trillions of gigabytes of critical data in Box: financial models, clinical trial protocols, M&amp;amp;A due diligence rooms, engineering schematics, and legal compliance playbooks. Up to this point, text-based search and retrieval-augmented generation (RAG) have successfully unlocked the vast narrative knowledge within these repositories, establishing a powerful and highly effective baseline for enterprise AI intelligence.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Traditional RAG architectures have mastered text processing, but the agentic era demands more. The next logical evolution is to extend this framework to capture the&lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt; &lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;inherently multimodal, deeply spatial, and highly structured elements that exist alongside text. While text embeddings excel at indexing prose, multimodal architectures unlock a major new capability: For example, they preserve the strict row-column semantics of financial tables, interpret visual evidence like clinical data, and map the logic of multi-page flowcharts without losing their spatial layout.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;To deliver next-generation capabilities that can handle the vast universe of digital content, Google Cloud and Box are &lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;integrating advanced multimodal capabilities into Box's Agentic Platform&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;, powered by &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/gemini-enterprise-agent-platform/models/gemini/embedding-2"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Gemini Multimodal Embeddings 2&lt;/span&gt;&lt;/a&gt;&lt;strong style="vertical-align: baseline;"&gt; &lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;merging Box’s industry-leading Intelligent Content Management platform with Google Cloud’s advanced AI embeddings.&lt;/span&gt;&lt;/p&gt;
&lt;h2&gt;&lt;strong style="vertical-align: baseline;"&gt;Benefits of improved embedding: Extending the dimensions of document content&lt;/strong&gt;&lt;/h2&gt;
&lt;ol&gt;
&lt;li aria-level="1" style="list-style-type: decimal; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Preserving visual and spatial geometry&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;: Complex document elements like multi-column tables or financial matrices rely on their spatial layout to convey meaning. Converting these elements into a flat string of text can disassociate column headers from their corresponding data points. Multimodal embeddings allow systems to interpret the document exactly as a human does, maintaining the integrity of spatial relationships.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: decimal; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Illuminating the visual modality&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;: Enterprise documents are filled with visual indicators: technical charts, process flowcharts, branding assets, and product photography. Multimodal capabilities ensure that these elements are no longer invisible to search systems, allowing users to query images and text simultaneously.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: decimal; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Connecting hybrid file formats&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;: Real-world business workflows rarely live in a single document format. An agent may need to cross-reference a PDF policy, a spreadsheet tracking log, and a presentation deck. Extending RAG with multimodal embeddings creates a unified understanding across these varied formats.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ol&gt;
&lt;h2&gt;&lt;strong style="vertical-align: baseline;"&gt;The Architectural Solution: Gemini Multimodal Embeddings 2&lt;/strong&gt;&lt;/h2&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Google Cloud’s &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/gemini-enterprise-agent-platform/models/gemini/embedding-2"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Gemini Multimodal Embeddings 2&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; introduces a unified, multimodal vector space capable of embedding text, raster images, document pages, rendered spreadsheet tables, and visual charts into the same semantic representation space.&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-image_full_width"&gt;






  
    &lt;div class="article-module h-c-page"&gt;
      &lt;div class="h-c-grid"&gt;
  

    &lt;figure class="article-image--large
      
      
        h-c-grid__col
        h-c-grid__col--6 h-c-grid__col--offset-3
        
        
      "
      &gt;

      
      
        
        &lt;img
            src="https://storage.googleapis.com/gweb-cloudblog-publish/original_images/image_bu28HMu.gif"
        
          alt="GIF_1"&gt;
        
        &lt;/a&gt;
      
    &lt;/figure&gt;

  
      &lt;/div&gt;
    &lt;/div&gt;
  




&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;Key product capabilities unlocked by gemini-embeddings-2:&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Crossmodal retrieval (text-to-visual / visual-to-text)&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;: Enables natural language queries to retrieve highly specific visual components, such as locating a target chart or diagram within a massive library of slides, without requiring manual tagging.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Layout-aware document embedding&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;: Rather than breaking files into arbitrary text blocks, the system can embed document page renderings directly, preserving visual hierarchies, callout boxes, and structural context.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Heterogeneous format bridging&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;: Native support for seamlessly bridging content across .docx, .xlsx, .pdf, .pptx, .png, and .csv without losing modality-specific structural information.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h2&gt;&lt;strong style="vertical-align: baseline;"&gt;Three core patterns of multimodal enterprise agents&lt;/strong&gt;&lt;/h2&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;By leveraging multimodal embeddings within Box, we have identified three uniqueprimary design patterns that illustrate how organizations can extend traditional RAG to support complex, visual workflows.&lt;/span&gt;&lt;/p&gt;
&lt;h3&gt;&lt;strong style="vertical-align: baseline;"&gt;Pattern 1: Complex financial &amp;amp; analytical reporting&lt;/strong&gt;&lt;/h3&gt;
&lt;h4&gt;&lt;strong style="vertical-align: baseline;"&gt;The challenge&lt;/strong&gt;&lt;/h4&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Corporate finance, research, and audit teams analyze highly structured documents where vital data resides in embedded tables, growth charts, and footnote annotations. Text-only indexing can separate these numbers from their context, making automated analysis challenging.&lt;/span&gt;&lt;/p&gt;
&lt;h4&gt;&lt;strong style="vertical-align: baseline;"&gt;The multimodal advantage&lt;/strong&gt;&lt;/h4&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Structural alignment&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;: The embedding model captures the physical structure of tables and charts, allowing financial agents to understand that a column header applies to a specific row of metrics.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Visual trend analysis&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;: Agents can cross-reference written summaries with visual trends in accompanying bar or line charts, identifying and pointing out discrepancies between written claims and source data.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Contextual sourcing&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;: Users can query complex portfolios and instantly retrieve the exact page, table, or chart supporting a specific metric.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;&lt;/div&gt;
&lt;div class="block-image_full_width"&gt;






  
    &lt;div class="article-module h-c-page"&gt;
      &lt;div class="h-c-grid"&gt;
  

    &lt;figure class="article-image--large
      
      
        h-c-grid__col
        h-c-grid__col--6 h-c-grid__col--offset-3
        
        
      "
      &gt;

      
      
        
        &lt;img
            src="https://storage.googleapis.com/gweb-cloudblog-publish/images/2_9rTykxw.max-1000x1000.png"
        
          alt="2"&gt;
        
        &lt;/a&gt;
      
    &lt;/figure&gt;

  
      &lt;/div&gt;
    &lt;/div&gt;
  




&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;h3&gt;&lt;strong style="vertical-align: baseline;"&gt;Pattern 2: Multimodal clinical decision support &amp;amp; assisted diagnosis&lt;/strong&gt;&lt;/h3&gt;
&lt;h4&gt;&lt;strong style="vertical-align: baseline;"&gt;The challenge&lt;/strong&gt;&lt;/h4&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;In healthcare and clinical environments, critical patient data is fragmented across vastly different, unstructured visual and textual formats — ranging from external physical photos (visual evidence) and microscopic pathology slides (lab reports) to structured risk matrices (triage grids). Traditional text-based systems or isolated analysis tools cannot synthesize these cross-modal relationships simultaneously, which can delay critical diagnoses or risk missing immediate, life-threatening procedural complications.&lt;/span&gt;&lt;/p&gt;
&lt;h4&gt;&lt;strong style="vertical-align: baseline;"&gt;The multimodal advantage&lt;/strong&gt;&lt;/h4&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Cross-modal clinical synthesis&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;: Evaluates physical symptoms alongside cellular-level laboratory evidence simultaneously by indexing clinical photos, histopathology imagery, and triage grids into a single space.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Granular anomaly identification&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;: Connects niche visual patterns under a microscope (like parasitic cyst walls) with medical knowledge to rapidly isolate rare conditions.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Risk-aware decision support&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;: Cross-references findings against triage frameworks to deliver instant warnings about immediate patient risks, such as life-threatening anaphylactic shock.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;&lt;/div&gt;
&lt;div class="block-image_full_width"&gt;






  
    &lt;div class="article-module h-c-page"&gt;
      &lt;div class="h-c-grid"&gt;
  

    &lt;figure class="article-image--large
      
      
        h-c-grid__col
        h-c-grid__col--6 h-c-grid__col--offset-3
        
        
      "
      &gt;

      
      
        
        &lt;img
            src="https://storage.googleapis.com/gweb-cloudblog-publish/images/3_ZPNWwdP.max-1000x1000.png"
        
          alt="3"&gt;
        
        &lt;/a&gt;
      
    &lt;/figure&gt;

  
      &lt;/div&gt;
    &lt;/div&gt;
  




&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;h3&gt;&lt;strong style="vertical-align: baseline;"&gt;Pattern 3: Cross-document multimodal synthesis &amp;amp; data reconciliation&lt;/strong&gt;&lt;/h3&gt;
&lt;h4&gt;&lt;strong style="vertical-align: baseline;"&gt;The challenge&lt;/strong&gt;&lt;/h4&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Enterprise information is fragmented across disconnected files and formats (e.g., PDF minutes, Excel charts, PNG flyers, and email threads). Traditional tools analyze these files in isolation, failing to connect the dots when verifying details or resolving data contradictions across independent documents.&lt;/span&gt;&lt;/p&gt;
&lt;h4&gt;&lt;strong style="vertical-align: baseline;"&gt;The multimodal advantage&lt;/strong&gt;&lt;/h4&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Cross-file synthesis&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;: Connects information across entirely different formats (PDFs, spreadsheets, images, emails) simultaneously to answer complex business queries.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Conflict resolution&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;: Flags and resolves contradictions between assets, such as catching outdated pricing on an image by cross-checking it against the latest financial spreadsheets.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Visual-to-text auditing&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;: Audits visual or scanned files against text-based records (e.g., verifying a signed PDF contract against a legal review email) to catch missing clauses or changes.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;&lt;/div&gt;
&lt;div class="block-image_full_width"&gt;






  
    &lt;div class="article-module h-c-page"&gt;
      &lt;div class="h-c-grid"&gt;
  

    &lt;figure class="article-image--large
      
      
        h-c-grid__col
        h-c-grid__col--6 h-c-grid__col--offset-3
        
        
      "
      &gt;

      
      
        
        &lt;img
            src="https://storage.googleapis.com/gweb-cloudblog-publish/images/4_IAwu96l.max-1000x1000.png"
        
          alt="4"&gt;
        
        &lt;/a&gt;
      
    &lt;/figure&gt;

  
      &lt;/div&gt;
    &lt;/div&gt;
  




&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;h2&gt;&lt;strong style="vertical-align: baseline;"&gt;The future of agentic enterprise content management&lt;/strong&gt;&lt;/h2&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;The integration of gemini-embeddings-2 into Box’s Agentic Platform is an important new capability to improve the next era of content intelligence. Multimodal embeddings help Box to move beyond basic search to active, intelligent collaboration.&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;Box's Intelligent Content Management platform represents a fundamental shift in enterprise AI infrastructure — moving beyond passive document storage to deliver a governed, semantically indexed reasoning layer where AI agents can interrogate, cross-reference, and act on content with full compliance and security controls already in place. &lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Powered by multimodal embeddings and a suite of native AI agents spanning search, metadata extraction, research, analysis, and composition, Box enables organizations to proactively surface insights such as flagging stale pricing data, expiring contract clauses, or cross-document contradictions before they become business risks. For high-complexity industries like financial services, life sciences, and legal operations, Box's ability to reason across text, tables, charts, and images makes multimodal understanding a competitive requirement. &lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Designed to interoperate with the broader enterprise AI ecosystem, Box serves as the single governed content foundation that ensures every AI-driven workflow is grounded in authorized, auditable enterprise data.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;When you think about it, the enterprise data landscape was always multimodal. Now we have the technology to make the most of it. By integrating gemini-embeddings-2, Box helps its users unlock unprecedented value from unstructured enterprise content. Product leaders who embrace multimodal-first architectures, rigorous precision benchmarking, and audit-ready grounding will lead the next wave of enterprise productivity and innovation.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="font-style: italic; vertical-align: baseline;"&gt;The team would like to thank Ken Ikeda, Afshaan Mazagonwalla, and Samip Thakkar for their work on this project.&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;</description><pubDate>Tue, 18 Aug 2026 16:00:00 +0000</pubDate><guid>https://cloud.google.com/blog/topics/partners/box-ai-agents-gemini-embeddings-multimodal-enterprise-ai/</guid><category>AI &amp; Machine Learning</category><category>Customers</category><category>Data Analytics</category><category>Partners</category><media:content height="540" url="https://storage.googleapis.com/gweb-cloudblog-publish/images/box-multimodal-agents-gemini-embeddings-head.max-600x600.png" width="540"></media:content><og xmlns:og="http://ogp.me/ns#"><type>article</type><title>How Box is unlocking multimodal enterprise agents with Gemini Embeddings 2</title><description></description><image>https://storage.googleapis.com/gweb-cloudblog-publish/images/box-multimodal-agents-gemini-embeddings-head.max-600x600.png</image><site_name>Google</site_name><url>https://cloud.google.com/blog/topics/partners/box-ai-agents-gemini-embeddings-multimodal-enterprise-ai/</url></og><author xmlns:author="http://www.w3.org/2005/Atom"><name>Sandhya Patil</name><title>Agentic Product Consulting Lead, Google</title><department></department><company></company></author><author xmlns:author="http://www.w3.org/2005/Atom"><name>Darryl Sladden</name><title>Staff AI Product Manager, Box</title><department></department><company></company></author></item><item><title>Building cost-effective, high-throughput gen AI workflows in Google Dataflow</title><link>https://cloud.google.com/blog/products/data-analytics/cost-effective-genai-workflows-in-google-dataflow/</link><description>&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Real-time streaming pipelines are the operational backbone of modern enterprises, continuously processing everything from customer support interactions to transaction logs. Traditionally, streaming DAGs are static; once deployed, their processing logic and execution paths are fixed. However, by integrating generative AI agents, we can move beyond static logic to adaptive execution. This allows streaming workflows to dynamically construct plans, query databases, and trigger custom remediation paths at runtime depending on the content of the data.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;For example, when a customer sends an angry message about a damaged order, a pipeline shouldn't just log the error or flag a dashboard. It should look up the order in the database that holds customer order and inventory records, decide on a remediation action (like shipping a replacement or issuing a refund), email the customer, and log the final resolution.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;However, streaming systems face a fundamental engineering hurdle when executing gen AI workflows: scale, latency, and cost. Sending every raw event directly to a heavyweight model or multi-step agent equipped with external database and email tools is prohibitively expensive, introduces high latency, and quickly exhausts API rate limits.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;This pattern addresses the scale and complexity challenge by combining &lt;/span&gt;&lt;a href="https://cloud.google.com/products/dataflow"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Google Dataflow&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;, Google Cloud's fully managed, serverless execution service for &lt;/span&gt;&lt;a href="https://beam.apache.org/" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Apache Beam&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;, and the &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/gemini-enterprise-agent-platform/build/adk"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Agent Development Kit&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; (ADK) to build a hybrid streaming pipeline. By using a lightweight, CPU-bound machine learning model upstream to filter and qualify events, we keep the pipeline highly cost-effective, routing only the complex cases to the downstream agent. There, the agent dynamically decides what actions to take, introducing dynamic branching to the stream without hardcoding thousands of conditional steps into the pipeline's static DAG.&lt;/span&gt;&lt;/p&gt;
&lt;h3&gt;&lt;strong style="vertical-align: baseline;"&gt;A universal blueprint for high-volume streams&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;While we use a customer support triage scenario below, this pre-filter + agentic action pattern is a universal paradigm. It applies to any stream where a high volume (&amp;gt;9X%) of events are routine, and only a small number require complex, contextual reasoning.&lt;/span&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;span style="vertical-align: baseline;"&gt;IT Operations &amp;amp; DevOps: Filtering millions of routine system logs on CPU, and triggering an agent to run diagnostics and open bug tickets only when a critical anomaly is flagged.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;span style="vertical-align: baseline;"&gt;Financial Fraud Triaging: Passing millions of transactions through lightweight, local rules, and calling an agent to execute multi-database lookup tools only for highly suspicious patterns.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;span style="vertical-align: baseline;"&gt;Industrial IoT: Monitoring normal telemetry on the edge, and routing erratic spikes to an agent to coordinate equipment shutdowns and email field engineers.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h2&gt;&lt;strong style="vertical-align: baseline;"&gt;The architecture: Why pre-filter streaming events?&lt;/strong&gt;&lt;/h2&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;In a high-throughput stream, the vast majority of messages do not require complex reasoning or remediation. They might be positive feedback, neutral inquiries, or simple queries.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Routing every single event to a heavyweight LLM workflow creates three primary bottlenecks:&lt;/span&gt;&lt;/p&gt;
&lt;ol&gt;
&lt;li aria-level="1" style="list-style-type: decimal; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;API cost:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Frontier models charge per token. Under high throughput, cost scales linearly with stream volume.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: decimal; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Latency:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Multi-step workflows (which involve database lookups and external API calls) take seconds, creating a bottleneck in streaming DAGs.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: decimal; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Quotas:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; External APIs have strict rate limits that streaming workers can easily exhaust.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;To prevent this, we build a pre-filtered pipeline in Apache Beam/Dataflow:&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-image_full_width"&gt;






  
    &lt;div class="article-module h-c-page"&gt;
      &lt;div class="h-c-grid"&gt;
  

    &lt;figure class="article-image--large
      
      
        h-c-grid__col
        h-c-grid__col--6 h-c-grid__col--offset-3
        
        
      "
      &gt;

      
      
        
        &lt;img
            src="https://storage.googleapis.com/gweb-cloudblog-publish/images/image1_XYX8VCT.max-1000x1000.jpg"
        
          alt="image1"&gt;
        
        &lt;/a&gt;
      
    &lt;/figure&gt;

  
      &lt;/div&gt;
    &lt;/div&gt;
  




&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;Pipeline flow&lt;/span&gt;&lt;/h3&gt;
&lt;ol&gt;
&lt;li aria-level="1" style="list-style-type: decimal; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Ingestion:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Read raw customer messages from &lt;/span&gt;&lt;a href="https://cloud.google.com/pubsub"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Google Pub/Sub&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: decimal; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Lightweight sentiment classifier (CPU):&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Run all messages through a lightweight, CPU-based Hugging Face model (&lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;distilbert-base-uncased-finetuned-sst-2-english&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt;) using Apache Beam’s &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;RunInference&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; transform. This executes locally on the Dataflow worker CPUs, avoiding external API costs.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: decimal; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Pre-qualification Gate:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; A simple &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;DoFn&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; filters the stream. Messages with &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;POSITIVE&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; or &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;NEUTRAL&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; sentiment are acknowledged and dropped.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: decimal; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Automated Remediation (ADK):&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; If and only if a message is classified as &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;NEGATIVE&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt;, we trigger the gen AI agent backed by &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;gemini-3.5-flash&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; using the &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;ADKAgentModelHandler&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt;. The agent uses tools to look up the user in &lt;/span&gt;&lt;a href="https://cloud.google.com/bigquery"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;BigQuery&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;, fetch orders, choose a remediation plan, and send a notification email via the Gmail API.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ol&gt;
&lt;h2&gt;&lt;strong style="vertical-align: baseline;"&gt;Adaptive execution: Making the Beam DAG dynamic&lt;/strong&gt;&lt;/h2&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;In traditional streaming architectures, the pipeline's Directed Acyclic Graph (DAG) is rigid. Once deployed to Dataflow, the sequence of transforms is set. If you need to handle new types of alerts or change how specific events are routed, you have to modify, test, and redeploy the entire pipeline.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;By placing a gen AI agent downstream of our sentiment pre-filter, we introduce a dynamic, adaptive node inside the static DAG.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;For the 95% of records that are positive or neutral, the pipeline runs along a fast, static path. But when the filter gates a negative record, the agent evaluates the payload and dynamically selects the correct sequence of API tools (e.g., database query, inventory check, or email notification) at runtime. This allows the pipeline to execute complex decision trees dynamically, eliminating the need to build and maintain thousands of hardcoded conditional branches in the static Apache Beam code.&lt;/span&gt;&lt;/p&gt;
&lt;h2&gt;&lt;strong style="vertical-align: baseline;"&gt;Implementing the pipeline&lt;/strong&gt;&lt;/h2&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Here is an example implementation in Apache Beam using the Google Agent Development Kit (ADK) and the &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;RunInference&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; framework.&lt;/span&gt;&lt;/p&gt;
&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;1. Defining the lightweight sentiment model&lt;/span&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;We define the upstream CPU model using &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;HuggingFacePipelineModelHandler&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt;. This model classifies sentiment into &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;POSITIVE&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt;, &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;NEUTRAL&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt;, or &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;NEGATIVE&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; on the worker instance.&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-code"&gt;&lt;dl&gt;
    &lt;dt&gt;code_block&lt;/dt&gt;
    &lt;dd&gt;&amp;lt;ListValue: [StructValue([(&amp;#x27;code&amp;#x27;, &amp;#x27;model_handler = HuggingFacePipelineModelHandler(\r\n    task=&amp;quot;sentiment-analysis&amp;quot;,\r\n    model=&amp;quot;distilbert-base-uncased-finetuned-sst-2-english&amp;quot;\r\n)&amp;#x27;), (&amp;#x27;language&amp;#x27;, &amp;#x27;&amp;#x27;), (&amp;#x27;caption&amp;#x27;, &amp;lt;wagtail.rich_text.RichText object at 0x7f18769ff460&amp;gt;)])]&amp;gt;&lt;/dd&gt;
&lt;/dl&gt;&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;2. Building the heavyweight ADK agent&lt;/span&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;The ADK agent acts as our remediation assistant. We equip it with three tools:&lt;/span&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;code style="vertical-align: baseline;"&gt;lookup_user&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt;: Queries BigQuery for the customer's email.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;code style="vertical-align: baseline;"&gt;lookup_orders&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt;: Queries BigQuery for the customer's orders and current product inventory.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;code style="vertical-align: baseline;"&gt;send_email&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt;: Sends a remediation email to the customer using the &lt;/span&gt;&lt;a href="https://developers.google.com/workspace/gmail/api/guides" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Gmail API&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;&lt;/div&gt;
&lt;div class="block-code"&gt;&lt;dl&gt;
    &lt;dt&gt;code_block&lt;/dt&gt;
    &lt;dd&gt;&amp;lt;ListValue: [StructValue([(&amp;#x27;code&amp;#x27;, &amp;#x27;def make_adk_tools(project: str, dataset: str = &amp;quot;sentiment_demo&amp;quot;):\r\n    def lookup_user(user_id: int) -&amp;gt; dict:\r\n        &amp;quot;&amp;quot;&amp;quot;Look up user information (email address) from BigQuery by user ID.&amp;quot;&amp;quot;&amp;quot;\r\n        from google.cloud import bigquery\r\n\r\n        client = bigquery.Client(project=project)\r\n        query = (\r\n            f&amp;quot;SELECT user_id, user_email &amp;quot;\r\n            f&amp;quot;FROM `{project}.{dataset}.users` &amp;quot;\r\n            f&amp;quot;WHERE user_id = @user_id&amp;quot;\r\n        )\r\n        job_config = bigquery.QueryJobConfig(\r\n            query_parameters=[bigquery.ScalarQueryParameter(&amp;quot;user_id&amp;quot;, &amp;quot;INT64&amp;quot;, user_id)]\r\n        )\r\n        try:\r\n            results = list(client.query(query, job_config=job_config).result())\r\n            if results:\r\n                row = results[0]\r\n                return {&amp;quot;user_id&amp;quot;: row.user_id, &amp;quot;user_email&amp;quot;: row.user_email}\r\n            return {&amp;quot;error&amp;quot;: f&amp;quot;No user found with user_id={user_id}&amp;quot;}\r\n        except Exception as exc:\r\n            return {&amp;quot;error&amp;quot;: str(exc)}\r\n\r\n    def lookup_orders(user_id: int) -&amp;gt; dict:\r\n        &amp;quot;&amp;quot;&amp;quot;Look up a user\&amp;#x27;s orders and current product inventory from BigQuery.&amp;quot;&amp;quot;&amp;quot;\r\n        from google.cloud import bigquery\r\n\r\n        client = bigquery.Client(project=project)\r\n        query = (\r\n            f&amp;quot;SELECT p.order_id, p.product_id, pr.remaining_inventory, pr.price &amp;quot;\r\n            f&amp;quot;FROM `{project}.{dataset}.purchases` p &amp;quot;\r\n            f&amp;quot;JOIN `{project}.{dataset}.products` pr ON p.product_id = pr.product_id &amp;quot;\r\n            f&amp;quot;WHERE p.user_id = @user_id&amp;quot;\r\n        )\r\n        job_config = bigquery.QueryJobConfig(\r\n            query_parameters=[bigquery.ScalarQueryParameter(&amp;quot;user_id&amp;quot;, &amp;quot;INT64&amp;quot;, user_id)]\r\n        )\r\n        try:\r\n            results = list(client.query(query, job_config=job_config).result())\r\n            orders = [\r\n                {\r\n                    &amp;quot;order_id&amp;quot;: row.order_id,\r\n                    &amp;quot;product_id&amp;quot;: row.product_id,\r\n                    &amp;quot;remaining_inventory&amp;quot;: row.remaining_inventory,\r\n                    &amp;quot;price&amp;quot;: float(row.price),\r\n                }\r\n                for row in results\r\n            ]\r\n            return {&amp;quot;orders&amp;quot;: orders}\r\n        except Exception as exc:\r\n            return {&amp;quot;error&amp;quot;: str(exc)}\r\n\r\n    def send_email(to_address: str, subject: str, body: str) -&amp;gt; str:\r\n        &amp;quot;&amp;quot;&amp;quot;Send a plain-text email to the customer via the Gmail API.&amp;quot;&amp;quot;&amp;quot;\r\n        import google.auth\r\n        import googleapiclient.discovery\r\n        import email.mime.text\r\n        import base64\r\n\r\n        try:\r\n            creds, _ = google.auth.default(\r\n                scopes=[&amp;quot;https://www.googleapis.com/auth/gmail.send&amp;quot;]\r\n            )\r\n            service = googleapiclient.discovery.build(&amp;quot;gmail&amp;quot;, &amp;quot;v1&amp;quot;, credentials=creds)\r\n\r\n            mime_msg = email.mime.text.MIMEText(body)\r\n            mime_msg[&amp;quot;to&amp;quot;] = to_address\r\n            mime_msg[&amp;quot;subject&amp;quot;] = subject\r\n            raw = base64.urlsafe_b64encode(mime_msg.as_bytes()).decode(&amp;quot;utf-8&amp;quot;)\r\n            service.users().messages().send(userId=&amp;quot;me&amp;quot;, body={&amp;quot;raw&amp;quot;: raw}).execute()\r\n            return &amp;quot;Email sent successfully&amp;quot;\r\n        except Exception as exc:\r\n            return f&amp;quot;Failed to send email: {exc}&amp;quot;\r\n\r\n    return [lookup_user, lookup_orders, send_email]&amp;#x27;), (&amp;#x27;language&amp;#x27;, &amp;#x27;&amp;#x27;), (&amp;#x27;caption&amp;#x27;, &amp;lt;wagtail.rich_text.RichText object at 0x7f18769ff430&amp;gt;)])]&amp;gt;&lt;/dd&gt;
&lt;/dl&gt;&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;We configure the &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;LlmAgent&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; and package it in the &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;ADKAgentModelHandler&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt;:&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-code"&gt;&lt;dl&gt;
    &lt;dt&gt;code_block&lt;/dt&gt;
    &lt;dd&gt;&amp;lt;ListValue: [StructValue([(&amp;#x27;code&amp;#x27;, &amp;#x27;adk_agent = LlmAgent(\r\n    name=&amp;quot;remediation_agent&amp;quot;,\r\n    model=&amp;quot;gemini-3.5-flash&amp;quot;,\r\n    instruction=(\r\n        &amp;quot;You are a customer service remediation assistant with access to &amp;quot;\r\n        &amp;quot;BigQuery lookup tools and an email sending tool. &amp;quot;\r\n        &amp;quot;When given a prompt describing a customer situation, follow the &amp;quot;\r\n        &amp;quot;numbered steps exactly and use your tools to complete the task.&amp;quot;\r\n    ),\r\n    tools=adk_tools,\r\n)\r\n\r\n# RunInference handler for the ADK agent\r\nadk_handler = ADKAgentModelHandler(agent=adk_agent)&amp;#x27;), (&amp;#x27;language&amp;#x27;, &amp;#x27;&amp;#x27;), (&amp;#x27;caption&amp;#x27;, &amp;lt;wagtail.rich_text.RichText object at 0x7f1877658340&amp;gt;)])]&amp;gt;&lt;/dd&gt;
&lt;/dl&gt;&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;3. Assembling the Dataflow DAG&lt;/span&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;The entire pipeline is declared cleanly. The upstream sentiment inference feeds directly into the filtering step (&lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;FilterNegativeADK&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt;), which then conditionally executes the downstream &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;ADKInference&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt;:&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-code"&gt;&lt;dl&gt;
    &lt;dt&gt;code_block&lt;/dt&gt;
    &lt;dd&gt;&amp;lt;ListValue: [StructValue([(&amp;#x27;code&amp;#x27;, &amp;#x27;with beam.Pipeline(options=pipeline_options) as p:\r\n    # 1. Read from Pub/Sub and classify sentiment on CPU\r\n    sentiment_results = (\r\n        p\r\n        | &amp;quot;ReadFromPubSub&amp;quot; &amp;gt;&amp;gt; beam.io.ReadFromPubSub(topic=known_args.input_topic)\r\n        | &amp;quot;DecodeMessages&amp;quot; &amp;gt;&amp;gt; beam.Map(lambda x: x.decode(\&amp;#x27;utf-8\&amp;#x27;))\r\n        | &amp;quot;SentimentInference&amp;quot; &amp;gt;&amp;gt; RunInference(model_handler)\r\n    )\r\n\r\n    # 2. Filter out non-negative sentiment and invoke the ADK Agent\r\n    _ = (\r\n        sentiment_results\r\n        | &amp;quot;FilterNegativeADK&amp;quot; &amp;gt;&amp;gt; beam.ParDo(FilterNegativeAndPromptADK())\r\n        | &amp;quot;ADKInference&amp;quot; &amp;gt;&amp;gt; RunInference(adk_handler)\r\n        | &amp;quot;LogADKResults&amp;quot; &amp;gt;&amp;gt; beam.ParDo(LogADKResponse())\r\n    )&amp;#x27;), (&amp;#x27;language&amp;#x27;, &amp;#x27;&amp;#x27;), (&amp;#x27;caption&amp;#x27;, &amp;lt;wagtail.rich_text.RichText object at 0x7f18772faac0&amp;gt;)])]&amp;gt;&lt;/dd&gt;
&lt;/dl&gt;&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;h2&gt;&lt;strong style="vertical-align: baseline;"&gt;Cost and performance advantages&lt;/strong&gt;&lt;/h2&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;By introducing this filtering step, we gain major engineering and operational advantages:&lt;/span&gt;&lt;/p&gt;
&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;1. Significant cost reductions&lt;/span&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Instead of paying for &lt;/span&gt;&lt;a href="https://ai.google.dev/gemini-api/docs/tokens" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Gemini input/output tokens&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; on 100% of incoming events, we pay only for the fraction that represent negative customer sentiment (typically &amp;lt; 5% of messages). The other 95% are classified locally on CPU instances at zero incremental API cost.&lt;/span&gt;&lt;/p&gt;
&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;2. High streaming throughput&lt;/span&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Dataflow distributes the CPU classification workload across many instances. Since CPU inference takes milliseconds, the pipeline scales horizontally to handle high-throughput event streams. The heavyweight LLM agent, which can take seconds per request due to tool execution, is called sparingly, preventing backlog.&lt;/span&gt;&lt;/p&gt;
&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;3. Native Apache Beam integration&lt;/span&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Adding the agent into the DAG requires no complex orchestration logic or manual thread pools. Using &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;ADKAgentModelHandler&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; with Beam's native &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;RunInference&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; transform handles parallel worker threads, batching, and integration automatically, keeping the codebase maintainable and clean.&lt;/span&gt;&lt;/p&gt;
&lt;h2&gt;&lt;strong style="vertical-align: baseline;"&gt;Key takeaways&lt;/strong&gt;&lt;/h2&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Streaming data is fast and high-volume, while heavyweight generative AI reasoning is slow and costly.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;By building a pre-filtered pipeline with Google Dataflow and the ADK, you get the best of both worlds: the cost and speed of local CPU-based models, and the deep, automated capabilities of Gemini-backed agents.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;To see the complete codebase and deploy this yourself, check out the &lt;/span&gt;&lt;a href="https://github.com/damccorm/next-2026-demo" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;next-2026-demo GitHub repository&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;.&lt;/span&gt;&lt;/p&gt;
&lt;hr/&gt;
&lt;p&gt;&lt;sub&gt;&lt;em&gt;&lt;span style="vertical-align: baseline;"&gt;Apache Beam is a trademark of the &lt;/span&gt;&lt;a href="https://www.apache.org/" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Apache Software Foundation&lt;/span&gt;&lt;/a&gt;&lt;/em&gt;&lt;/sub&gt;&lt;/p&gt;&lt;/div&gt;</description><pubDate>Tue, 18 Aug 2026 16:00:00 +0000</pubDate><guid>https://cloud.google.com/blog/products/data-analytics/cost-effective-genai-workflows-in-google-dataflow/</guid><category>AI &amp; Machine Learning</category><category>Streaming</category><category>Data Analytics</category><og xmlns:og="http://ogp.me/ns#"><type>article</type><title>Building cost-effective, high-throughput gen AI workflows in Google Dataflow</title><description></description><site_name>Google</site_name><url>https://cloud.google.com/blog/products/data-analytics/cost-effective-genai-workflows-in-google-dataflow/</url></og><author xmlns:author="http://www.w3.org/2005/Atom"><name>Reza Rokni</name><title>Group Product Manager</title><department></department><company></company></author><author xmlns:author="http://www.w3.org/2005/Atom"><name>Danny McCormick</name><title>Software Engineer, Google Cloud</title><department></department><company></company></author></item><item><title>Building operational resilience with agentic AI in financial services</title><link>https://cloud.google.com/blog/topics/financial-services/building-operational-resilience-with-agentic-ai-in-financial-services/</link><description>&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;span&gt;&lt;span style="vertical-align: baseline;"&gt;For financial institutions, operational resilience has long been embedded in regulatory and supervisory expectations — to say nothing of the high expectations of consumers. With the implementation of the European Union’s &lt;/span&gt;&lt;a href="https://cloud.google.com/blog/products/identity-security/the-eus-dora-has-arrived-google-cloud-is-ready-to-help"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Digital Operational Resiliency Act&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; (DORA), those expectations have become even more stringent, with more explicit, harmonized, and evidence-driven requirements. Firms must now demonstrate that their critical business services and supporting digital infrastructures can withstand disruption, support coordinated response, and recover with control.&lt;/span&gt;&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span&gt;&lt;span style="vertical-align: baseline;"&gt;To meet these conditions, &lt;/span&gt;&lt;a href="https://www.db.com/" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Deutsche Bank&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; developed an AI-powered agentic resilience platform that modernized its regulatory &lt;/span&gt;&lt;a href="https://cloud.google.com/transform/tabletopping-the-tabletop-new-perspectives-cybersecurity-favorite-role-playing-game"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;tabletop resilience exercises&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; at scale and turned manual preparation into context-aware and evidence-ready simulations grounded in actual operational data. The platform builds enterprise context from architecture, data flows, logs, incident history, alerting signals, and operational telemetry to generate scenarios, simulated operational evidence, structured session records, and regulator-ready artifacts.&lt;/span&gt;&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;At many large banks with operations that span interdependent applications, data flows, and third-party services, this is a critical and even existential shift. Across financial services, supervisory expectations are evolving and as they do, banks’ tabletop exercises must reflect their production dependencies, real operating conditions, and compliance with consistent evidence standards more directly.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;As Deutsche Bank considered how to successfully and efficiently make this shift at scale, it looked to &lt;/span&gt;&lt;a href="https://www.db.com/news/detail/20201204-deutsche-bank-and-google-cloud-sign-pioneering-cloud-and-innovation-partnership?language_id=1" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;its long-time partner, Google Cloud&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;, and its growing suite of agentic AI tools.&lt;/span&gt;&lt;/p&gt;
&lt;h3&gt;&lt;strong style="vertical-align: baseline;"&gt;From tabletop exercises to resilience intelligence&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span&gt;&lt;span style="vertical-align: baseline;"&gt;With its agentic resilience platform, DB has been able to transform its tabletop exercises from manual preparation to a continuous intelligence model. And it’s been able to extend the same agentic layer to root-cause analysis when real operational context is needed.&lt;/span&gt;&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;This means that every scenario it runs is based on real enterprise signals. The platform can then reflect true system dependencies, failure patterns, and business impact instead of relying on static inputs that are more likely to return assumptions than real-time insights.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;By using &lt;/span&gt;&lt;a href="https://cloud.google.com/products/gemini-enterprise-agent-platform"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Gemini Enterprise Agent Platform&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;, DB has been able to migrate this operational context into structured scenarios with clear timelines, decision points, and expected responses. This has ensured that each exercise is grounded in real system behavior that produces consistent, audit-ready evidence that meets regulatory expectations.&lt;/span&gt;&lt;/p&gt;
&lt;h3&gt;&lt;strong style="vertical-align: baseline;"&gt;Dual orchestration for control and flexibility&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;In order to deliver both regulator-grade control and operational flexibility, DB’s platform introduced a dual-orchestration architecture that separates workflows into two complementary execution models.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;First, for regulator-aligned execution, the bank is using &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/gemini-enterprise-agent-platform/scale/runtime/use-a-langgraph-agent"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;LangGraph&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; to ensure that it generates every scenario through a traceable, deterministic process — with clear lineage from input context to output — that supports the auditability required for supervisory review.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Next, for its adaptive and investigative scenarios, DB is using &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/gemini-enterprise-agent-platform/build/adk"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Google Agent Development Kit&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; (ADK) to enable agent-driven coordination. This approach allows the bank’s platform to dynamically analyze conditions and generate responses without predefined execution paths.&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-image_full_width"&gt;






  
    &lt;div class="article-module h-c-page"&gt;
      &lt;div class="h-c-grid"&gt;
  

    &lt;figure class="article-image--large
      
      
        h-c-grid__col
        h-c-grid__col--6 h-c-grid__col--offset-3
        
        
      "
      &gt;

      
      
        
        &lt;img
            src="https://storage.googleapis.com/gweb-cloudblog-publish/images/image1_dbc2MvR.max-1000x1000.png"
        
          alt="image1"&gt;
        
        &lt;/a&gt;
      
        &lt;figcaption class="article-image__caption "&gt;&lt;p data-block-key="jys9x"&gt;Figure 1. Architecture for context assembly, orchestration, and scenario generation.&lt;/p&gt;&lt;/figcaption&gt;
      
    &lt;/figure&gt;

  
      &lt;/div&gt;
    &lt;/div&gt;
  




&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;span&gt;&lt;span style="vertical-align: baseline;"&gt;With this architectural separation, the platform can combine governed execution with adaptive investigation while preserving a common intelligence layer. The same agents and tools can reason over architecture, data-flow diagrams, logs, and code artifacts across tabletop scenario generation and related incident-analysis workflows. Importantly, this supports a &lt;/span&gt;&lt;a href="https://cloud.google.com/blog/products/identity-security/cyber-snapshot-report-enterprise-resilience-key-to-toolchain-success"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;consistent resilience model&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; across both planned exercises and real operational events.&lt;/span&gt;&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="font-style: italic; vertical-align: baseline;"&gt;Deutsche Bank’s&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt; objective with this &lt;/span&gt;&lt;span style="font-style: italic; vertical-align: baseline;"&gt;platform&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt; was to engineer a resilience model for critical financial systems that meets regulatory expectations — even within highly complex, distributed environments. By linking dynamically generated scenarios to real business context and combining governed orchestration with adaptive analysis, the platform has given us an intelligent, continuously adaptive model for operational resilience.”&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt; – &lt;/span&gt;&lt;strong style="font-style: italic; vertical-align: baseline;"&gt;Sanjay Tripathi&lt;/strong&gt;&lt;span style="font-style: italic; vertical-align: baseline;"&gt;, Managing Director, Global Head of Surveillance Technology &amp;amp; Compliance Cloud &amp;amp; AI Transformation Lead, Deutsche Bank&lt;/span&gt;&lt;/p&gt;
&lt;h3&gt;&lt;strong style="vertical-align: baseline;"&gt;Powering generation and governance with Google Cloud&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Google Cloud’s suite of agentic tools is providing the foundation for scaling Deutsche Bank’s platform across its many governed, enterprise-grade resilience workflows. Here’s how:&lt;/span&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;a href="https://cloud.google.com/run"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Cloud Run&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; supports elastic execution of scenario and evidence-generation services. &lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;a href="https://cloud.google.com/products/gemini-enterprise-agent-platform"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Gemini Enterprise Agent Platform&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; transforms operational context into structured resilience scenarios.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;a href="https://docs.cloud.google.com/gemini-enterprise-agent-platform/build/adk"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Google ADK&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; enables adaptive agent coordination.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;a href="https://cloud.google.com/sql"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Cloud SQL&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; provides durable persistence for scenarios, session artifacts, and review records.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Collectively, these services give DB support for the traceable generation, controlled execution, and persistent evidence record required for compliance review and continuous improvement.&lt;/span&gt;&lt;/p&gt;
&lt;h3&gt;&lt;strong style="vertical-align: baseline;"&gt;Scalable, evidence-ready resilience testing&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Every scenario generated by Deutsche Bank’s platform drives a structured tabletop session for the teams that run response, escalation, and recovery. Because these exercises are grounded in real enterprise context, they reflect operational reality while also strengthening consistency across teams and creating audit-ready evidence that meets regulatory expectations. For institutions that operate under DORA or similar frameworks, this makes it easier to demonstrate controlled, coordinated, and disciplined response at scale. &lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;This model is now being applied across multiple DB portfolios, which is helping the bank establish more consistent and scalable resilience paradigms and a replicable blueprint for the broader financial sector.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;In this model, root-cause analysis acts as the feedback loop between real incidents and future resilience testing. The resulting insights from production events can inform future tabletop scenarios, while exercise outcomes can strengthen response playbooks, escalation paths, and recovery readiness.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;All of this extends the platform’s value from planned resilience exercises to real operational events while keeping scenario-based resilience testing as the primary use case.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;As adoption expands, this platform brings consistency by embedding Google Cloud’s methodology for context-aware resilience. It eliminates fragmented manual approaches and establishes a cross-functional, AI-informed operating model across the bank.&lt;/span&gt;&lt;/p&gt;
&lt;h3&gt;&lt;strong style="vertical-align: baseline;"&gt;Toward resilience intelligence&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;The bank’s next step is to extend this approach into a broader resilience intelligence layer, which is possible because it can deploy the same patterns to support playbook refinement, recovery-readiness assessments, and continuous validation of controls against evolving system conditions.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;For financial institutions, this is a strategic shift. As systems become more distributed and regulatory expectations more demanding, banks must move from periodic resilience testing to continuous, intelligence-driven capabilities. At Deutsche Bank, Google Cloud is making that transition simple across the organization.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="font-style: italic; vertical-align: baseline;"&gt;Learn more about Google Cloud’s methodology for context-aware resilience in this &lt;/span&gt;&lt;a href="https://cloud.google.com/blog/topics/financial-services/improve-financial-resilience-with-google-cloud?e=0"&gt;&lt;span style="font-style: italic; text-decoration: underline; vertical-align: baseline;"&gt;article&lt;/span&gt;&lt;/a&gt;&lt;span style="font-style: italic; vertical-align: baseline;"&gt;.&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;</description><pubDate>Tue, 18 Aug 2026 14:00:00 +0000</pubDate><guid>https://cloud.google.com/blog/topics/financial-services/building-operational-resilience-with-agentic-ai-in-financial-services/</guid><category>AI &amp; Machine Learning</category><category>Customers</category><category>Financial Services</category><media:content height="540" url="https://storage.googleapis.com/gweb-cloudblog-publish/images/deutsche-bank-operational-resilience-agentic.max-600x600.png" width="540"></media:content><og xmlns:og="http://ogp.me/ns#"><type>article</type><title>Building operational resilience with agentic AI in financial services</title><description></description><image>https://storage.googleapis.com/gweb-cloudblog-publish/images/deutsche-bank-operational-resilience-agentic.max-600x600.png</image><site_name>Google</site_name><url>https://cloud.google.com/blog/topics/financial-services/building-operational-resilience-with-agentic-ai-in-financial-services/</url></og><author xmlns:author="http://www.w3.org/2005/Atom"><name>Pankaj Ojha</name><title>Director &amp; Lead Architect – Agentic Resilience Platform, Deutsche Bank</title><department></department><company></company></author><author xmlns:author="http://www.w3.org/2005/Atom"><name>Florian Graf</name><title>Staff Solutions Consultant, Google Cloud Consulting</title><department></department><company></company></author></item><item><title>Using BigQuery Graphs with measures for trusted agentic workloads</title><link>https://cloud.google.com/blog/products/data-analytics/bigquery-graphs-with-measures-for-trusted-agentic-workloads/</link><description>&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;When enterprises transition from using simple chat assistants to autonomous, agentic workloads, they quickly run into a hard truth: Agents are prone to inaccurate insights when working with directly raw tables. &lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;a href="https://docs.cloud.google.com/bigquery/docs/graph-measures"&gt;&lt;strong style="text-decoration: underline; vertical-align: baseline;"&gt;BigQuery Graph&lt;/strong&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; helps organizations move beyond flat, static tables to represent enterprises exactly how they exist in the physical world: as interconnected business entities with real-world dependencies. With the support of measures in BigQuery Graph (preview), we are unifying governed metrics with relationship mapping. This allows your agents to reason across complex dependencies captured in graphs with precision of measures.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Why relationships matter&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Traditional data structures are blind to multi-hop business context, causing AI agents to make incorrect operational decisions:&lt;/span&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;The concrete problem:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; If a retailer has an agent who is asked why winter jacket sales dropped 12% in Seattle, it can query flat tables to report the &lt;/span&gt;&lt;span style="font-style: italic; vertical-align: baseline;"&gt;what&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt; (the 12% dip). But it fails at the &lt;/span&gt;&lt;span style="font-style: italic; vertical-align: baseline;"&gt;why&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt; because it cannot trace the relational path: &lt;/span&gt;&lt;span style="font-style: italic; vertical-align: baseline;"&gt;Seattle orders&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt; ➔ &lt;/span&gt;&lt;span style="font-style: italic; vertical-align: baseline;"&gt;distribution centers&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt; ➔ &lt;/span&gt;&lt;span style="font-style: italic; vertical-align: baseline;"&gt;suppliers delayed by regional storms&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;The risk of disjointed systems:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Lacking relationship context, the agent suggests an irrelevant 15% markdown campaign, needlessly eroding margins. Furthermore, maintaining separate systems - where one team maps supplier relationships in a separate graph database while another maintains SQL metrics - forces your agent to stitch these stacks together at runtime. This process is slow, expensive, and leads to inconsistent KPI calculations.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Measures in BigQuery Graph solves this by letting you &lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;map existing tables to a property graph in-place with zero ETL&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;. This unified setup enables a logical evolution of inquiry:&lt;/span&gt;&lt;/p&gt;
&lt;ol&gt;
&lt;li aria-level="1" style="list-style-type: decimal; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Metadata grounding&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; establishes &lt;/span&gt;&lt;span style="font-style: italic; vertical-align: baseline;"&gt;what&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt; data you have.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: decimal; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Business metrics (measures)&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; calculate &lt;/span&gt;&lt;span style="font-style: italic; vertical-align: baseline;"&gt;how&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt; your business performed.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: decimal; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Relationship mapping (graph)&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; uncovers &lt;/span&gt;&lt;span style="font-style: italic; vertical-align: baseline;"&gt;why&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt; it happened.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ol&gt;
&lt;h3&gt;&lt;strong style="vertical-align: baseline;"&gt;Under the hood&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Historically, standard SQL joins during graph traversals duplicate rows, leading to incorrect aggregation calculations. BigQuery Graph solves this natively.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Data modelers define a &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;MEASURE&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; (like &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;SUM&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; or &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;AVG&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt;) directly within the Property Graph DDL. Using standard SQL via the &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;GRAPH_EXPAND&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; function and the &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;AGG&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; aggregator, the engine resolves the structural graph paths &lt;/span&gt;&lt;span style="font-style: italic; vertical-align: baseline;"&gt;before&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt; evaluating metrics. This ensures your agent is smart enough to know when it needs a &lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;calculator (SQL)&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; and when it needs a &lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;map (graph)&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Because public projects like &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;bigquery-public-data&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; are strictly read-only, you must map the logical property graph inside your own project using a placeholder variable (&lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;YOUR_PROJECT_ID&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt;), while directly referencing the read-only public tables as nodes and edges.&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-code"&gt;&lt;dl&gt;
    &lt;dt&gt;code_block&lt;/dt&gt;
    &lt;dd&gt;&amp;lt;ListValue: [StructValue([(&amp;#x27;code&amp;#x27;, &amp;#x27;-- 1. Map the graph inside YOUR project \r\n\r\n\r\nCREATE OR REPLACE PROPERTY GRAPH `YOUR_PROJECT_ID.YOUR_DATASET.thelook_ecommerce_graph`\r\nNODE TABLES(\r\n  `bigquery-public-data.thelook_ecommerce.users` AS User\r\n    KEY(id)\r\n    LABEL User PROPERTIES(id, city, country),\r\n  `bigquery-public-data.thelook_ecommerce.orders` AS Order\r\n    KEY(order_id)\r\n    LABEL Order PROPERTIES(\r\n      order_id, \r\n      MEASURE(AVG(num_of_item)) AS avg_items_per_order,\r\n      MEASURE(SUM(num_of_item)) AS total_items\r\n    )\r\n)\r\nEDGE TABLES(\r\n  `bigquery-public-data.thelook_ecommerce.orders` AS OrderedBy\r\n    SOURCE KEY(order_id) REFERENCES Order(order_id)\r\n    DESTINATION KEY(user_id) REFERENCES User(id)\r\n    LABEL ORDERED_BY\r\n);\r\n\r\n-- 2. Query your new graph with standard SQL—using standard {Label}_{Property} column outputs\r\nSELECT\r\n  User_city AS city,\r\n  ROUND(AGG(Order_avg_items_per_order), 2) AS agg_avg_items,\r\n  ROUND(AGG(Order_total_items), 2) AS agg_total_items\r\nFROM GRAPH_EXPAND(&amp;quot;YOUR_PROJECT_ID.YOUR_DATASET.thelook_ecommerce_graph&amp;quot;)\r\nGROUP BY User_city\r\nORDER BY agg_total_items DESC\r\nLIMIT 10;&amp;#x27;), (&amp;#x27;language&amp;#x27;, &amp;#x27;&amp;#x27;), (&amp;#x27;caption&amp;#x27;, &amp;lt;wagtail.rich_text.RichText object at 0x7f1876bd6c70&amp;gt;)])]&amp;gt;&lt;/dd&gt;
&lt;/dl&gt;&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;h3&gt;&lt;strong style="vertical-align: baseline;"&gt;Democratizing graph intelligence in BigQuery Studio&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;To make managing and deploying these relationship networks frictionless for both developers and business users, we have built native, intuitive operational tools directly into BigQuery Studio:&lt;/span&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Visual graph modeler:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; A no-code, drag-and-drop interface inside BigQuery Studio that lets you visually build, edit, and map property graphs, nodes, and edges without writing complex DDL scripts manually.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;&lt;/div&gt;
&lt;div class="block-image_full_width"&gt;






  
    &lt;div class="article-module h-c-page"&gt;
      &lt;div class="h-c-grid"&gt;
  

    &lt;figure class="article-image--large
      
      
        h-c-grid__col
        h-c-grid__col--6 h-c-grid__col--offset-3
        
        
      "
      &gt;

      
      
        
        &lt;img
            src="https://storage.googleapis.com/gweb-cloudblog-publish/original_images/1_CXQhslw.gif"
        
          alt="1"&gt;
        
        &lt;/a&gt;
      
    &lt;/figure&gt;

  
      &lt;/div&gt;
    &lt;/div&gt;
  




&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;ul&gt;
&lt;li&gt;&lt;strong style="vertical-align: baseline;"&gt;Conversational Analytics (CA) integration:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Users can interact with the graph naturally. Instead of guessing table joins, Conversational Analytics agents navigate the deterministic, relationship-aware map of the graph, converting natural language questions into precise, boundary-constrained GoogleSQL or ISO GQL queries. This prevents model hallucinations and enforces semantic consistency.&lt;/span&gt;&lt;/li&gt;
&lt;/ul&gt;&lt;/div&gt;
&lt;div class="block-image_full_width"&gt;






  
    &lt;div class="article-module h-c-page"&gt;
      &lt;div class="h-c-grid"&gt;
  

    &lt;figure class="article-image--large
      
      
        h-c-grid__col
        h-c-grid__col--6 h-c-grid__col--offset-3
        
        
      "
      &gt;

      
      
        
        &lt;img
            src="https://storage.googleapis.com/gweb-cloudblog-publish/original_images/2_23oI53E.gif"
        
          alt="2"&gt;
        
        &lt;/a&gt;
      
    &lt;/figure&gt;

  
      &lt;/div&gt;
    &lt;/div&gt;
  




&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;h3&gt;&lt;strong style="vertical-align: baseline;"&gt;Unified semantics: Native Looker integration&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;To avoid maintaining fragmented logic stacks, business metrics must live at the data layer. By integrating &lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;Looker (LookML)&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; natively with BigQuery Graphs as&lt;/span&gt;&lt;a href="https://docs.cloud.google.com/looker/docs/analytic-models"&gt;&lt;span style="vertical-align: baseline;"&gt; &lt;/span&gt;&lt;strong style="text-decoration: underline; vertical-align: baseline;"&gt;in-database analytic models&lt;/strong&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;, you define logic once at the core:&lt;/span&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong style="vertical-align: baseline;"&gt;Database-managed models (sql_analytic_model_name):&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Point Looker directly to your database-defined BigQuery Graph using &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;sql_analytic_model_name&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; to map standard LookML dimensions and measures directly to your graph properties.&lt;/span&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong style="vertical-align: baseline;"&gt;Looker-managed models (derived_analytic_model):&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Define your BigQuery Graph schema directly inside your LookML view using &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;derived_analytic_model&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt;. Looker will dynamically generate and execute the SQL DDL statements to maintain the graph inside BigQuery.&lt;/span&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong style="vertical-align: baseline;"&gt;Enterprise DevOps workflows:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Manage your graph's entire lifecycle using the &lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;Looker IDE, Git-based version control, and Continuous Integration (CI)&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;. Core KPIs (like Churn Rate) remain completely identical, verified, and trusted.&lt;/span&gt;&lt;/li&gt;
&lt;/ul&gt;&lt;/div&gt;</description><pubDate>Thu, 13 Aug 2026 17:00:00 +0000</pubDate><guid>https://cloud.google.com/blog/products/data-analytics/bigquery-graphs-with-measures-for-trusted-agentic-workloads/</guid><category>AI &amp; Machine Learning</category><category>Databases</category><category>Data Analytics</category><og xmlns:og="http://ogp.me/ns#"><type>article</type><title>Using BigQuery Graphs with measures for trusted agentic workloads</title><description></description><site_name>Google</site_name><url>https://cloud.google.com/blog/products/data-analytics/bigquery-graphs-with-measures-for-trusted-agentic-workloads/</url></og><author xmlns:author="http://www.w3.org/2005/Atom"><name>Deepak Dayama</name><title>Group Product Manager, Google Cloud</title><department></department><company></company></author><author xmlns:author="http://www.w3.org/2005/Atom"><name>Yun Zhang</name><title>Software Development Manager</title><department></department><company></company></author></item><item><title>Looker’s semantic layer governs Gemini Enterprise data for user trust</title><link>https://cloud.google.com/blog/products/business-intelligence/integrating-looker-and-gemini-enterprise/</link><description>&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;For organizations deploying AI agents at scale, there’s often a critical divide between structured and unstructured data. While large language models (LLMs) excel at parsing text documents, emails, and PDFs, they can struggle when presented with raw enterprise databases. Meanwhile, standard natural-language-to-SQL (NL2SQL) models often guess how database schemas fit together, which can lead to unpredictable queries, inconsistent metrics, and AI hallucinations that erode user trust. &lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;a href="https://cloud.google.com/gemini-enterprise?utm_source=google&amp;amp;utm_medium=cpc&amp;amp;utm_campaign=1713704-Workspace-DR-APAC-IN-en-Google-BKWS-MIX-Hybrid-GeminiEnterprise&amp;amp;utm_content=c-Hybrid+%7C+BKWS+-+EXA+%7C+Txt-Gemini+Enterprise-Generic-435278751514&amp;amp;utm_term=gemini+enterprise&amp;amp;gclsrc=aw.ds&amp;amp;gad_source=1&amp;amp;gad_campaignid=23381004221&amp;amp;gclid=Cj0KCQjwlqTRBhCBARIsANrkrxgdte2Ry_iXPSiT0N7_AEygV0lPiuKDnLLm9uPdxiQKE39HhLfoQsgaAgdmEALw_wcB&amp;amp;e=0"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Gemini Enterprise&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; brings the best of Google AI to every employee through an intuitive chat interface that acts as a single front door for AI in the workplace. And now, Looker’s governed semantic layer serves as the trusted foundation for structured data within Gemini Enterprise, enabling trusted self-service business intelligence for all Gemini Enterprise users. With this integration, Looker analysts and admins can publish conversational agents natively into their Gemini Enterprise environments via the Agent-to-Agent (A2A) protocol. Now, organizations can provide their AI-accelerated taskforce with robust and trusted tools, powered by real-time analytics, that they can explore in natural language in addition to their daily workspace workflows. Making it easy to offer conversational agents in Gemini Enterprise expands discoverability and promotes a data-driven culture, while reducing friction to adoption.&lt;/span&gt;&lt;/p&gt;
&lt;h3&gt;&lt;strong style="vertical-align: baseline;"&gt;Bringing a semantic foundation to structured and unstructured data&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;By combining Looker’s semantic layer with Gemini Enterprise, you can query both structured databases and unstructured documents in plain English, all in one place. Instead of jumping between dashboards and other tools to understand your numbers, teams can instantly connect hard metrics with real-world context to solve problems and make decisions faster.&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-image_full_width"&gt;






  
    &lt;div class="article-module h-c-page"&gt;
      &lt;div class="h-c-grid"&gt;
  

    &lt;figure class="article-image--large
      
      
        h-c-grid__col
        h-c-grid__col--6 h-c-grid__col--offset-3
        
        
      "
      &gt;

      
      
        
        &lt;img
            src="https://storage.googleapis.com/gweb-cloudblog-publish/original_images/1_Ei9b2UE.gif"
        
          alt="1"&gt;
        
        &lt;/a&gt;
      
        &lt;figcaption class="article-image__caption "&gt;&lt;p data-block-key="wxusy"&gt;Publishing Looker agents for consumption in Gemini Enterprise&lt;/p&gt;&lt;/figcaption&gt;
      
    &lt;/figure&gt;

  
      &lt;/div&gt;
    &lt;/div&gt;
  




&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;h3&gt;&lt;strong style="vertical-align: baseline;"&gt;Minimize AI hallucinations&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;If you ask the typical AI chatbot to calculate "revenue" or "churn rate" against an unstructured cloud database, it has to guess which tables to join, which filters to apply, and which timestamps to trust. This can result in different people asking the same question, only to get completely different answers.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Looker’s semantic layer eliminates this guesswork, serving critical context to Gemini Enterprise in the form of codified data, allowing the agent to give deterministic, predictable responses.&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-code"&gt;&lt;dl&gt;
    &lt;dt&gt;code_block&lt;/dt&gt;
    &lt;dd&gt;&amp;lt;ListValue: [StructValue([(&amp;#x27;code&amp;#x27;, &amp;#x27;[ Gemini Enterprise Chat UI ] \r\n               │\r\n      (A2A Protocol / NLP)\r\n               ▼\r\n   [ Looker Governed Agent ]  ──► Generates Deterministic SQL\r\n               │\r\n   [ Looker Semantic Layer ]  ──► Business-Approved Definitions &amp;amp; Logic\r\n               │\r\n               ▼\r\n     [ Enterprise Data Cloud ]  ──► (BigQuery, AlloyDB, Spanner, etc.)&amp;#x27;), (&amp;#x27;language&amp;#x27;, &amp;#x27;&amp;#x27;), (&amp;#x27;caption&amp;#x27;, &amp;lt;wagtail.rich_text.RichText object at 0x7f1876bbe670&amp;gt;)])]&amp;gt;&lt;/dd&gt;
&lt;/dl&gt;&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;When a Gemini Enterprise user requests a business KPI in Gemini Enterprise, the request is routed directly to a Looker agent. The semantic layer generates deterministic, precise SQL based on version-controlled business logic. This helps ensure when an executive asks for "Revenue," they get the exact, governed enterprise metric — not a guess.&lt;/span&gt;&lt;/p&gt;
&lt;h3&gt;&lt;strong style="vertical-align: baseline;"&gt;Robust governance and secure access management&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Data governance and security are critical when introducing AI to enterprise data warehouses. Organizations can’t risk corporate information being loosely ingested, indexed, or exposed outside of strict permissions.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Looker’s integration with Gemini Enterprise is built on a zero-risk pass-through architecture, processing the data, but not writing to persistent storage. Gemini Enterprise does &lt;/span&gt;&lt;span style="font-style: italic; vertical-align: baseline;"&gt;not&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt; ingest, replicate, or store your underlying database records. Instead, the integration operates safely and securely over the A2A protocol, following these core tenets:&lt;/span&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;OAuth authorization:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; In order to interact with a Looker agent within Gemini Enterprise, end users provide a secure, one-time OAuth consent. This binds their Gemini session to their specific Looker credentials.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Strong governance enforcement:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Because the architecture relies on live pass-through queries, Looker’s existing row-level and column-level access controls are maintained.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Strict security isolation:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; If a user does not have permission to view, say, sensitive regional payroll or financial rows within the Looker platform, the Looker agent actively restricts that data in the Gemini environment. Should an agent be published to the Agent Gallery to simplify discovery, it still does not bypass the security controls that you established.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h3&gt;&lt;strong style="vertical-align: baseline;"&gt;Technical capabilities and enterprise readiness&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Deploying Looker agents natively into Gemini Enterprise via the &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/gemini/data-agents/conversational-analytics-api/integration-patterns#a2a-orchestration"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;A2A protocol&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; doesn't just make it smarter — it makes it more interactive and interoperable, without sacrificing security. Here are some of the features you’ll find in this release.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;Rich visual interactivity: support for charts&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;They say a picture is worth a thousand words. &lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;When users interact with Looker agents inside Gemini Enterprise, the platform goes beyond textual explanations and provides native, interactive data charts. If a user asks for a visual trend—such as monthly sales performance or regional distribution—the Looker agent maps the database response with rich, presentation-ready visualizations directly inside the universal chat box.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;Note:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; If you published Looker agents in Gemini Enterprise prior to Looker release 26.12, we recommend &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/looker/docs/conversational-analytics-looker-data-agents#republish-agent-ge"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;updating or refreshing&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; them to take advantage of these enhanced visualization capabilities.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;&lt;strong&gt;I&lt;/strong&gt;&lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;nteroperability with first- and third-party agents&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Looker agents published to Gemini Enterprise can understand context across different agents and data sources. Leveraging standard communication frameworks, these agents can securely share structured, governed insights with other first-party Google Cloud agents like the &lt;/span&gt;&lt;a href="https://ai.google.dev/gemini-api/docs/deep-research" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Deep Research Agent&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; or external third-party agents to create structured workflows. This enables complex multi-agent orchestration, where an enterprise operational agent can pull data from a Looker agent to feed into a separate productivity or supply-chain workflow.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;Looker-based user authentication&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;To preserve enterprise governance, this integration implements a robust, identity-centric authentication model. Users are required to provide a one-time OAuth consent, binding their active Gemini Enterprise session securely to their underlying Looker credentials. This helps ensure that every conversational query hitting your databases is authenticated at the user level, enforcing pre-existing Looker permission structures, row-level data access filters, and column-level masking rules — no exceptions.&lt;/span&gt;&lt;/p&gt;
&lt;h3&gt;&lt;strong style="vertical-align: baseline;"&gt;Trusted data in Gemini Enterprise&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;The future of work is agentic. Gemini Enterprise provides a single, secure architecture to deploy a global digital task force,empowering your business with the best of Google AI for developers, employees, and customers.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;The integration of Looker with Gemini Enterprise not only brings trusted data analytics to business users but also adds rich interactivity, visual charts, and data storytelling directly into their everyday workspace. As business users embrace this agentic new way of working, they aren't just getting text answers; they are getting presentation-ready visualizations that bring operational metrics to life and deliver complex insights. &lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;To get started, &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/looker/docs/conversational-analytics-looker-data-agents#publish-data-agents"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;learn how to publish your data agents in Gemini Enterprise&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; to make your agent’s predefined context and analytics available to your entire organization.&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;</description><pubDate>Tue, 11 Aug 2026 16:00:00 +0000</pubDate><guid>https://cloud.google.com/blog/products/business-intelligence/integrating-looker-and-gemini-enterprise/</guid><category>AI &amp; Machine Learning</category><category>Data Analytics</category><category>Business Intelligence</category><og xmlns:og="http://ogp.me/ns#"><type>article</type><title>Looker’s semantic layer governs Gemini Enterprise data for user trust</title><description></description><site_name>Google</site_name><url>https://cloud.google.com/blog/products/business-intelligence/integrating-looker-and-gemini-enterprise/</url></og><author xmlns:author="http://www.w3.org/2005/Atom"><name>Tarunima Tripathi</name><title>Product Manager</title><department></department><company></company></author></item><item><title>How WPP operationalizes platform and data engineering for AI marketing</title><link>https://cloud.google.com/blog/products/media-entertainment/how-wpp-operationalizes-platform-and-data-engineering-for-ai-marketing/</link><description>&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Between chaotic levels of market fragmentation and economic volatility, marketing and communications agencies can no longer rely on the human intuition they’ve traditionally used to win clients and optimize their ad spend. WPP is replacing that guesswork with an AI-powered view of shifting market dynamics, giving brands predictive certainty that lets them invest with confidence while moving at the speed of the market. That’s the value of &lt;/span&gt;&lt;a href="https://www.wpp.com/en/open" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;WPP Open&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;, its agentic marketing system.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;But before it could begin applying sophisticated AI models to power those insights, &lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;WPP had to overcome a critical engineering challenge: the marketing data that made up the models was fragmented across hundreds of global agencies. While this dynamic made it nearly impossible to deploy AI tools efficiently and securely, access to models was only part of the equation. And  until it built a reliable way to ingest, clean, and serve data to those models, WPP couldn’t unlock the true potential of generative AI.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;To solve this, WPP partnered with Google Cloud to construct a unified data backbone and  custom platform engineering path. Now, by standardizing its serverless compute patterns and data processing workflows, WPP is able to  securely deploy targeted marketing campaigns in days instead of months.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;Architecting a centralized, service-based data foundation &lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;An important part of this effort was accelerating data availability and centralizing management. To do this, WPP adopted a service-based project structure for its current production environment. Rather than isolating every workload into separate silos, its engineering team centralized &lt;/span&gt;&lt;a href="https://cloud.google.com/storage"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Google Cloud Storage&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; (GCS) and &lt;/span&gt;&lt;a href="https://cloud.google.com/bigquery"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;BigQuery&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; into dedicated, shared data projects, while also segregating the compute and processing workloads into distinct processing projects.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;This structure simplified the core team’s user experience and ensured that all data consumers interacted with a unified source of truth. Because data from WPP’s various product lines lives in shared infrastructure, it was essential that security be strictly enforced at a granular level. By directly applying identity and access management (IAM) controls at the individual GCS bucket and BigQuery dataset levels, the company’s teams only see the data they’re  authorized to access.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;At the same time, raw data from various partners lands in dedicated GCS buckets in order to keep the raw inputs organized and isolated. From there, &lt;/span&gt;&lt;a href="https://cloud.google.com/products/managed-service-for-apache-spark"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Managed Service for Apache Spark&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; executes custom Apache Spark jobs to cleanse, normalize, and canonicalize information into standardized cohort definitions (SCDs). By utilizing a serverless architecture combined with &lt;/span&gt;&lt;a href="https://www.kubeflow.org/" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Kubeflow&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; for pipeline orchestration, WPP’s data engineering team avoided the overhead that often results from managing cluster infrastructure. This allowed them to focus entirely on the data transformation logic fueling the downstream GCS and BigQuery layers  that ultimately feed the company’s audience &amp;amp; performance AI models.&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-image_full_width"&gt;






  
    &lt;div class="article-module h-c-page"&gt;
      &lt;div class="h-c-grid"&gt;
  

    &lt;figure class="article-image--large
      
      
        h-c-grid__col
        h-c-grid__col--6 h-c-grid__col--offset-3
        
        
      "
      &gt;

      
      
        
        &lt;img
            src="https://storage.googleapis.com/gweb-cloudblog-publish/original_images/wpp_data_flow_architecture.jpg"
        
          alt="1 WPP Data Pipeline Architecture"&gt;
        
        &lt;/a&gt;
      
    &lt;/figure&gt;

  
      &lt;/div&gt;
    &lt;/div&gt;
  




&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;p style="padding-left: 40px;"&gt;&lt;code style="font-style: italic; vertical-align: baseline;"&gt;What made our collaboration with Google Cloud successful was the balance they struck between uncompromising professionalism when it comes to best practices and timely delivery of incredibly pragmatic, real-world solutions.&lt;br/&gt;&lt;/code&gt;&lt;code style="font-style: italic; vertical-align: baseline;"&gt;- Jonas Dahlbaek&lt;br/&gt;&lt;/code&gt;&lt;code style="font-style: italic; vertical-align: baseline;"&gt;Senior Data Engineering Lead, WPP&lt;/code&gt;&lt;/p&gt;
&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;Standardizing data into unified cohorts&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt; &lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;When raw data enters WPP’s processing zone, its platform converts it into SCDs that become core concepts used throughout the framework for keying purposes. These are based on five keys: age, gender, geo, product, and interest. But these underlying data definitions are fluid and continuously canonicalized to reflect evolving marketing concepts. As a result, this uniform structure allows WPP to join and aggregate data on a global scale without exposing sensitive underlying particulars or relying on shared identifiers.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;The platform's core processing engine was built in type-safe Scala to ensure comprehensive visibility and compliance This custom framework tightly controls how data is transformed, and it inherently supports full source traceability while guaranteeing that every data point within the curated datasets can be traced back to its origin. This is a crucial level of traceability when building enterprise AI applications, as data scientists and auditors must understand exactly what information feeds into the models, even as WPP concurrently prepares to transition to Google Cloud &lt;/span&gt;&lt;a href="https://cloud.google.com/products/knowledge-catalog"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Knowledge Catalog&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; for automated, enterprise-wide data governance in the future.&lt;/span&gt;&lt;/p&gt;
&lt;p style="padding-left: 40px;"&gt;&lt;code style="font-style: italic; vertical-align: baseline;"&gt;Working with Google Cloud has been instrumental in accelerating and standardizing our engineering efforts. In a world where massive volumes of fragmented data present a daily challenge, having the right infrastructure is paramount to thriving in the AI age and helps our developers and AI marketers alike.&lt;/code&gt;&lt;br/&gt;&lt;code style="font-style: italic; vertical-align: baseline;"&gt;-Suleman Khan&lt;br/&gt;&lt;/code&gt;&lt;code style="font-style: italic; vertical-align: baseline;"&gt;Product Manager for OI &amp;amp; Google Partnerships, WPP&lt;/code&gt;&lt;/p&gt;
&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;Standardizing the enterprise software lifecycle&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;For WPP, even with all these steps in place, processing data is only half the battle. To serve applications and manage the underlying infrastructure, the company’s platform engineering team developed a suite of reusable and centralized GitLab continuous integration and continuous deployment (CI/CD) templates. With this, WPP reduced the cognitive load on individual development teams and ensured that all deployments met strict corporate security standards.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;These templates manage various enterprise workloads autonomously. The suite includes universal &lt;/span&gt;&lt;a href="https://cloud.google.com/run"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Cloud Run&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; templates for full-stack web applications and  batch data processing and scheduled pipelines. It also includes a deploy-only template for multi-stage workflows and a &lt;/span&gt;&lt;a href="https://cloud.google.com/functions"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Cloud Run functions&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; deployment template for event-driven microservices.&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-image_full_width"&gt;






  
    &lt;div class="article-module h-c-page"&gt;
      &lt;div class="h-c-grid"&gt;
  

    &lt;figure class="article-image--large
      
      
        h-c-grid__col
        h-c-grid__col--6 h-c-grid__col--offset-3
        
        
      "
      &gt;

      
      
        
        &lt;img
            src="https://storage.googleapis.com/gweb-cloudblog-publish/images/2_WPP_Cloud_Platform_Engineering.max-1000x1000.png"
        
          alt="2 WPP Cloud Platform Engineering"&gt;
        
        &lt;/a&gt;
      
    &lt;/figure&gt;

  
      &lt;/div&gt;
    &lt;/div&gt;
  




&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;Implementing zero-rebuild promotion&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Rebuilding container images in a production environment can introduce unnecessary risk and the potential for configuration drift. In order to maintain environmental consistency, WPP embraced a "build once, deploy many" methodology that applied cross-project IAM logic and &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/artifact-registry/docs"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Google Cloud Artifact Registry&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; configurations.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;As part of this process, developers build and test container images in the development environment. Once those exact, immutable container images are validated, they’re promote  directly to production. This zero-rebuild promotion ensures total parity across deployment stages and eliminates unexpected production behaviors. The CI/CD templates also facilitate progressive traffic migration, which allowed teams to route a small percentage of traffic to new revisions before initiating a full rollout.&lt;/span&gt;&lt;/p&gt;
&lt;p style="padding-left: 40px;"&gt;&lt;code style="font-style: italic; vertical-align: baseline;"&gt;Immutable deployments. Traceable data. Unshakable trust. When you know exactly what goes into your AI, you can ship at the speed of light.&lt;/code&gt;&lt;br/&gt;&lt;code style="font-style: italic; vertical-align: baseline;"&gt;- Ranjith K Poldas&lt;br/&gt;&lt;/code&gt;&lt;code style="font-style: italic; vertical-align: baseline;"&gt;Associate Director , Devops (I&amp;amp;P), WPP Media&lt;/code&gt;&lt;/p&gt;
&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;Automating security and intelligent networking&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;With this modern architecture, enterprise security acts as a foundational enabler for WPP, so it integrated &lt;/span&gt;&lt;a href="https://cloud.google.com/wiz"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Wiz&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; security scanning directly into the pre-push phase of the CI/CD pipeline to catch vulnerabilities before code merges. The company also utilized &lt;/span&gt;&lt;a href="https://cloud.google.com/security/products/iap"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Google Cloud Identity-Aware Proxy&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; to enforce zero-trust access across its  internal applications.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;To further simplify operations, WPP adopted templates with intelligent virtual private cloud (VPC) logic. This configuration automatically identifies and resolves networking conflicts between legacy VPC connectors and modern &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/run/docs/configuring/vpc-direct-vpc"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Direct VPC&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; access. This automated networking prevents deployment failures and accelerates the release cycle.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;Monitoring operational health and driving ROI&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Because a resilient platform foundation requires deep observability, WPP’s engineering team now monitors strict operational metrics instead of relying solely on deployment frequency. The team tracks request latency across p50, p95, and p99 percentiles, alongside 4xx and 5xx error rates. It  also monitors container startup times to mitigate cold starts, while tracking overall CPU and memory utilization. This granularity ensures that both data pipelines and serverless infrastructure always remain highly available.&lt;/span&gt;&lt;/p&gt;
&lt;p style="padding-left: 40px;"&gt;&lt;code style="font-style: italic; vertical-align: baseline;"&gt;"Navigating a transformation of this scale across multiple complex workstreams—spanning data engineering, platform infrastructure, and AI integration—required more than just alignment; it demanded deep, mutual trust. Working as true partners, Google Cloud and WPP moved in lockstep to deliver production-ready platform capabilities on time."&lt;br/&gt;&lt;/code&gt;&lt;code style="font-style: italic; vertical-align: baseline;"&gt;Yang Yue , Program Manager , Google Cloud&lt;/code&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;For WPP, operationalizing its data and AI stacks at this velocity provided the necessary infrastructure for its advanced workloads, and the business impact was clear and quantifiable. By building this dual foundation, the company reduced creative and strategy time from four weeks to just three hours. It also saw a 70% gain in production efficiency, a 33x increase in content volume, and  a 2.8x increase in campaign return on investment. In short, by partnering with Google Cloud and implementing a broad suite of products and tools, WPP was able to quickly realize a significant ROI and boost productivity, efficiency, reliability, and security across the company.&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;</description><pubDate>Mon, 10 Aug 2026 16:00:00 +0000</pubDate><guid>https://cloud.google.com/blog/products/media-entertainment/how-wpp-operationalizes-platform-and-data-engineering-for-ai-marketing/</guid><category>AI &amp; Machine Learning</category><category>Customers</category><category>Media &amp; Entertainment</category><media:content height="540" url="https://storage.googleapis.com/gweb-cloudblog-publish/images/wpp-ai-platform-engineering.max-600x600.png" width="540"></media:content><og xmlns:og="http://ogp.me/ns#"><type>article</type><title>How WPP operationalizes platform and data engineering for AI marketing</title><description></description><image>https://storage.googleapis.com/gweb-cloudblog-publish/images/wpp-ai-platform-engineering.max-600x600.png</image><site_name>Google</site_name><url>https://cloud.google.com/blog/products/media-entertainment/how-wpp-operationalizes-platform-and-data-engineering-for-ai-marketing/</url></og><author xmlns:author="http://www.w3.org/2005/Atom"><name>Utkarsh Bhardwaj</name><title>Technical Solutions Consultant</title><department></department><company></company></author><author xmlns:author="http://www.w3.org/2005/Atom"><name>Prabha Arya</name><title>Strategic Cloud Engineer</title><department></department><company></company></author></item><item><title>How Malachyte solves retail’s cold-start problem with managed real-time AI</title><link>https://cloud.google.com/blog/products/data-analytics/solving-retails-cold-start-problem-malachytes-recommendation-reinvention/</link><description>&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;What’s the best way to recommend products to little-known users? &lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;We’ve spent our careers trying to solve this problem for major companies like Spotify and Priceline, and it’s why Sidd founded &lt;/span&gt;&lt;a href="https://www.malachyte.com/" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Malachyte&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;, an AI-powered ecommerce recommendation platform. These days, consumers have come to expect content that feels personalized and relevant, and online services competing for their attention have no choice but to do this exceptionally well.  &lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Malachyte was inspired by some unique insights into how advanced AI models, and large language models in particular, could be applied in new ways to old challenges like personalization and recommendations. &lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;As Malachyte set out to win potential customers’ business, we needed secure, scalable, reliable and, above all, leading-edge AI infrastructure to continue building the personalization algorithm we had always envisioned. By utilizing Google Cloud tools like &lt;/span&gt;&lt;a href="https://cloud.google.com/bigtable"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Bigtable&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; and &lt;/span&gt;&lt;a href="https://cloud.google.com/products/managed-service-for-apache-kafka"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Managed Service for Apache Kafka&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;, Malachyte has been able to help some of its retailers &lt;/span&gt;&lt;a href="https://www.malachyte.com/case-studies" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;double and sometimes even triple&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; their sales. &lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;This is the story of how we built it, and the ways any founder can use services like these to start deploying AI foundation models in new ways.&lt;/span&gt;&lt;/p&gt;
&lt;h3&gt;&lt;strong style="vertical-align: baseline;"&gt;How Malachyte lifted sales for their users &lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;For Malachyte, the aha moment was discovering that it could use neural networks with attention mechanisms — the same concept powering large language models — to personalize retail search and product pages. This approach is what enables LLMs to derive meaning from the relative order of items in a sequence, in their case the order of words and syllables in a sentence. When it comes to a retail website or app, what Malachyte wanted to capture was the sequence of customer interactions with the site.&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-image_full_width"&gt;






  
    &lt;div class="article-module h-c-page"&gt;
      &lt;div class="h-c-grid"&gt;
  

    &lt;figure class="article-image--large
      
      
        h-c-grid__col
        h-c-grid__col--6 h-c-grid__col--offset-3
        
        
      "
      &gt;

      
      
        
        &lt;img
            src="https://storage.googleapis.com/gweb-cloudblog-publish/images/1_-_Malachyte_blog_.max-1000x1000.png"
        
          alt="1 - Malachyte blog"&gt;
        
        &lt;/a&gt;
      
        &lt;figcaption class="article-image__caption "&gt;&lt;p data-block-key="o24ff"&gt;What if we predicted the next thing a user wants on an ecommerce website just like LLMs predict the next word in a sentence?&lt;/p&gt;&lt;/figcaption&gt;
      
    &lt;/figure&gt;

  
      &lt;/div&gt;
    &lt;/div&gt;
  




&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;A pre-GPT language model might have tried to look at a specific sequence of words or even fragments of words (what we now know of as tokens), but those earlier models wouldn’t examine what happens if the words were in the comparable order but weren’t contiguous or were re-arranged. The breakthrough came — in part through Google’s work on transformers — when LLMs gained the ability to understand complex and long-range dependencies within a sequence of items. &lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;This more sophisticated method has delivered dramatic results — both for the proliferation of gen AI in general, and for Malachyte’s application of the technology.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;To make this work in practice, Malachyte creates a vector of everything known about a visitor when they arrive on a site.  Most users are visiting for the first time, so little is known about them. This is what’s known as the  “cold start” problem. The trick is to use every interaction with a user to refine this vector. Each new addition to the vector, like a click or a query, does two things: it drives a prediction about the next thing the user wants, and it provides more information about the user.  &lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Malachyte’s platform then updates the user vector and the prediction at the same time. This not only enhances the understanding of the individual user and their preferences, it also improves the overall model with the anonymized user data. With every inference, the context of both the average and the specific shopper grows. &lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;The company further innovates by not just using attention-based neural networks but combining that with updating user profiles 100 milliseconds at time. &lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-image_full_width"&gt;






  
    &lt;div class="article-module h-c-page"&gt;
      &lt;div class="h-c-grid"&gt;
  

    &lt;figure class="article-image--large
      
      
        h-c-grid__col
        h-c-grid__col--6 h-c-grid__col--offset-3
        
        
      "
      &gt;

      
      
        
        &lt;img
            src="https://storage.googleapis.com/gweb-cloudblog-publish/images/2_-_Malachyte_Blog.max-1000x1000.png"
        
          alt="2 - Malachyte Blog"&gt;
        
        &lt;/a&gt;
      
        &lt;figcaption class="article-image__caption "&gt;&lt;p data-block-key="o24ff"&gt;Malachyte’s recommendation and search agents populate the next page’s search results or recommendation carousels based on what users clicked on previous pages.&lt;/p&gt;&lt;/figcaption&gt;
      
    &lt;/figure&gt;

  
      &lt;/div&gt;
    &lt;/div&gt;
  




&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;To be sure, this idea isn’t in itself new. Retailers have long used collaborative filtering recommendations systems to identify similar users and items that required massive sets of interaction history. These models typically required a lot of data, including third-party cookie-based profiles and demographics. &lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;By focusing on the sequence of interactions in a session, retailers can achieve far more personalization — with less required data or spend — than by focusing only on a user’s profile. As a bonus, retailers can now offer their users more privacy by not relying on long-term cookie data.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;This works because of the model structure and multimodal vectors that encode everything they know about a user, including browser data, click history and searches. The output, too, is multimodal: The same model can be applied to on-site search product pages, category pages, and add-to-cart carousels.  &lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;To make this work, each product in the catalog is embedded into the same space as the user vector, which gets updated and subsequently moves the vector closer to relevant products and further from those that aren’t. The neural network computing the embedding is being continuously trained across retailers who work with Malachyte, improving the quality for everyone. The &lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;system effectively becomes a data cooperative with each retailer's user helping make the model smarter for everyone.&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-image_full_width"&gt;






  
    &lt;div class="article-module h-c-page"&gt;
      &lt;div class="h-c-grid"&gt;
  

    &lt;figure class="article-image--large
      
      
        h-c-grid__col
        h-c-grid__col--6 h-c-grid__col--offset-3
        
        
      "
      &gt;

      
      
        
        &lt;img
            src="https://storage.googleapis.com/gweb-cloudblog-publish/images/3_-_Malachyte_blog_vector_space_-_high_res.max-1000x1000.png"
        
          alt="3 - Malachyte blog vector space - high res"&gt;
        
        &lt;/a&gt;
      
        &lt;figcaption class="article-image__caption "&gt;&lt;p data-block-key="o24ff"&gt;A user session represented as a vector in a space of products.&lt;/p&gt;&lt;/figcaption&gt;
      
    &lt;/figure&gt;

  
      &lt;/div&gt;
    &lt;/div&gt;
  




&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;To make this delivery for every user at every inference in 100 milliseconds, Malachite found real benefits in building onGoogle Cloud’s real-time AI stack.   &lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;With this system, every behavioral event streams into a Managed Service for Apache Kafka cluster. Rather than queuing for a future training job, each event immediately becomes an update to the user’s profile in Bigtable. The Kafka cluster allows the customer’s front-end to persist, so the user session signals quickly with little worry about how they fit into the user vector.  &lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Bigtable allows Malachyte’s services to look up and update the right user vectors, and it and Kafka operate at the order of 10 milliseconds per step, which allows the entire recommendation loop to complete with no disruption to the user experience.  &lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-image_full_width"&gt;






  
    &lt;div class="article-module h-c-page"&gt;
      &lt;div class="h-c-grid"&gt;
  

    &lt;figure class="article-image--large
      
      
        h-c-grid__col
        h-c-grid__col--6 h-c-grid__col--offset-3
        
        
      "
      &gt;

      
      
        
        &lt;img
            src="https://storage.googleapis.com/gweb-cloudblog-publish/images/4_-_Malachyte_Blog.max-1000x1000.png"
        
          alt="4 - Malachyte Blog"&gt;
        
        &lt;/a&gt;
      
        &lt;figcaption class="article-image__caption "&gt;&lt;p data-block-key="o24ff"&gt;The three layer real-time AI architecture: a retailer’s website, Malachyte’s AI models and serving front ends, and context management infrastructure.&lt;/p&gt;&lt;/figcaption&gt;
      
    &lt;/figure&gt;

  
      &lt;/div&gt;
    &lt;/div&gt;
  




&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;In addition to a fast core, a second layer of product catalog updates, inventory signals, and retailer-specific dimensional data keeps product data up to date. This flows through &lt;/span&gt;&lt;a href="https://cloud.google.com/pubsub"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Cloud Pub/Sub&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;, which offers globally accessible REST APIs that enable connections retailers can use without deep integration work. Malachyte agents run on &lt;/span&gt;&lt;a href="https://cloud.google.com/kubernetes-engine"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Google Kubernetes Engine &lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;(GKE), with model inference on &lt;/span&gt;&lt;a href="https://cloud.google.com/products/compute"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Google Compute Engine&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; (GCE).&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-image_full_width"&gt;






  
    &lt;div class="article-module h-c-page"&gt;
      &lt;div class="h-c-grid"&gt;
  

    &lt;figure class="article-image--large
      
      
        h-c-grid__col
        h-c-grid__col--6 h-c-grid__col--offset-3
        
        
      "
      &gt;

      
      
        
        &lt;img
            src="https://storage.googleapis.com/gweb-cloudblog-publish/images/5_Malachyte_blog.max-1000x1000.png"
        
          alt="5 Malachyte blog"&gt;
        
        &lt;/a&gt;
      
        &lt;figcaption class="article-image__caption "&gt;&lt;p data-block-key="o24ff"&gt;Continuous ingestion of external data, such as product catalog updates, operates through Pub/Sub’s global messaging system.&lt;/p&gt;&lt;/figcaption&gt;
      
    &lt;/figure&gt;

  
      &lt;/div&gt;
    &lt;/div&gt;
  




&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;With its migration to Google Cloud’s AI architecture, Malachyte demonstrated that production AI inference and training are about more than GPUs and storage. They require real-time continuous learning infrastructure that includes a fast key-value store, a streaming layer, and a managed messaging system, all integrated with the foundation model architecture. &lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;This approach also shows that even a small team like Malachyte’s can have a big impact in an industry. It just needs access to powerful infrastructure and core AI managed services.&lt;/span&gt;&lt;/p&gt;
&lt;h3&gt;&lt;strong style="vertical-align: baseline;"&gt;Try it for yourself &lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Looking to shake up your industry or stay ahead of the competition like Malachyte? Try &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/managed-service-for-apache-kafka/docs/overview"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Managed Service for Apache Kafka&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;, &lt;/span&gt;&lt;a href="https://cloud.google.com/pubsub"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Cloud Pub/Sub&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;, and &lt;/span&gt;&lt;a href="https://cloud.google.com/bigtable"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Bigtable&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;. New customers can receive &lt;/span&gt;&lt;a href="https://cloud.google.com/free"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;$300 in Google Cloud credits&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;. &lt;/span&gt;&lt;/p&gt;&lt;/div&gt;</description><pubDate>Mon, 10 Aug 2026 16:00:00 +0000</pubDate><guid>https://cloud.google.com/blog/products/data-analytics/solving-retails-cold-start-problem-malachytes-recommendation-reinvention/</guid><category>AI &amp; Machine Learning</category><category>Customers</category><category>Retail</category><category>Data Analytics</category><media:content height="540" url="https://storage.googleapis.com/gweb-cloudblog-publish/images/malachyte-ai-foundation-models-retail-recomm.max-600x600.png" width="540"></media:content><og xmlns:og="http://ogp.me/ns#"><type>article</type><title>How Malachyte solves retail’s cold-start problem with managed real-time AI</title><description></description><image>https://storage.googleapis.com/gweb-cloudblog-publish/images/malachyte-ai-foundation-models-retail-recomm.max-600x600.png</image><site_name>Google</site_name><url>https://cloud.google.com/blog/products/data-analytics/solving-retails-cold-start-problem-malachytes-recommendation-reinvention/</url></og><author xmlns:author="http://www.w3.org/2005/Atom"><name>Sidd Motwani</name><title>CEO, Malachyte</title><department></department><company></company></author><author xmlns:author="http://www.w3.org/2005/Atom"><name>Vicki Boykis</name><title>Staff Machine Learning Engineer, Malachyte</title><department></department><company></company></author></item><item><title>Google named a Leader in The Forrester Wave™: AI Platforms, Q3 2026</title><link>https://cloud.google.com/blog/products/ai-machine-learning/google-named-a-leader-in-the-forrester-wave-ai-platforms/</link><description>&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;At Google Cloud, we help organizations of all sizes build and operationalize complex agentic workflows with total confidence. By combining world-class AI research with an open, fully integrated AI platform, we give customers the flexibility to innovate and the foundation to deliver measurable business value. &lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;At the center of it all is Gemini Enterprise, a unified platform designed to &lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;power the agentic enterprise&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;, &lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;meet builders where they are&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;, and &lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;deliver enterprise trust by default&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;. &lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;We believe this integrated approach is why Google has been named a &lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;Leader&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; in &lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;The Forrester Wave&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;™:&lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt; AI Platforms, Q3 2026 report&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;, and received the &lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;highest score in the Strategy category&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;. &lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-image_full_width"&gt;






  
    &lt;div class="article-module h-c-page"&gt;
      &lt;div class="h-c-grid"&gt;
  

    &lt;figure class="article-image--large
      
      
        h-c-grid__col
        h-c-grid__col--6 h-c-grid__col--offset-3
        
        
      "
      &gt;

      
      
        
        &lt;img
            src="https://storage.googleapis.com/gweb-cloudblog-publish/images/Image_AI-Platforms-Q3-2026.max-1000x1000.png"
        
          alt="Image_AI-Platforms,-Q3-2026"&gt;
        
        &lt;/a&gt;
      
    &lt;/figure&gt;

  
      &lt;/div&gt;
    &lt;/div&gt;
  




&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;Powering the Agentic Era&lt;/span&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;As agents become embedded across every part of the business, organizations need a unified platform. Gemini Enterprise serves as the front door to AI for your entire organization, removing the silos between business users, developers, and IT leaders by connecting them all with shared, universal context.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Gemini Enterprise Agent Platform is the foundation that enables technical teams and IT leaders to safely, securely, and cost-effectively build and deploy production-grade agents. Any agent built in Agent Platform can be deployed across your entire workforce through the Gemini Enterprise app, putting custom agentic capabilities directly into the hands of every employee. &lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;By unifying enterprise data, frontier model capabilities, developer tooling, and IT operations under one roof, Gemini Enterprise empowers teams to transform products, services, and complex agentic tasks while maintaining centralized control every step of the way.&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-video"&gt;



&lt;div class="article-module article-video "&gt;
  &lt;figure&gt;
    &lt;a class="h-c-video h-c-video--marquee"
      href="https://youtube.com/watch?v=NfHXzDXgUBs"
      data-glue-modal-trigger="uni-modal-NfHXzDXgUBs-"
      data-glue-modal-disabled-on-mobile="true"&gt;

      
        

        &lt;div class="article-video__aspect-image"
          style="background-image: url(https://storage.googleapis.com/gweb-cloudblog-publish/images/image1_JF2Osf0.max-1000x1000.png);"&gt;
          &lt;span class="h-u-visually-hidden"&gt;New Way Now: Mars accelerates marketing campaigns from months to weeks with Gemini Enterprise&lt;/span&gt;
        &lt;/div&gt;
      
      &lt;svg role="img" class="h-c-video__play h-c-icon h-c-icon--color-white"&gt;
        &lt;use xlink:href="#mi-youtube-icon"&gt;&lt;/use&gt;
      &lt;/svg&gt;
    &lt;/a&gt;

    
      &lt;figcaption class="article-video__caption h-c-page"&gt;
        
          &lt;h4 class="h-c-headline h-c-headline--four h-u-font-weight-medium h-u-mt-std"&gt;See how Mars accelerates their marketing campaigns from months to weeks with Gemini Enterprise&lt;/h4&gt;
        
        
      &lt;/figcaption&gt;
    
  &lt;/figure&gt;
&lt;/div&gt;

&lt;div class="h-c-modal--video"
     data-glue-modal="uni-modal-NfHXzDXgUBs-"
     data-glue-modal-close-label="Close Dialog"&gt;
   &lt;a class="glue-yt-video"
      data-glue-yt-video-autoplay="true"
      data-glue-yt-video-height="99%"
      data-glue-yt-video-vid="NfHXzDXgUBs"
      data-glue-yt-video-width="100%"
      href="https://youtube.com/watch?v=NfHXzDXgUBs"
      ng-cloak&gt;
   &lt;/a&gt;
&lt;/div&gt;

&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;h3&gt;&lt;strong style="vertical-align: baseline;"&gt;Meeting builders where they are&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;No two development teams build agents in the same way. Some work in coding environments, others rely on low-code tools, and many organizations use a mix of models. Gemini Enterprise is built for that reality. We offer a platform tailored to every team's skill set, enabling high-code developers to build complex agentic systems while giving business and operational teams low-cod and no-code tools to rapidly design and test agent behaviors. And with access to over 200 native and third-party models, alongside pre-built agent templates, engineering teams can move from prototype to deployment with speed. &lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;True enterprise AI should extend far beyond text. Gemini Enterprise is multi-modal by design, meaning it natively understands, reasons across, and generates text, code, audio, image, and video inputs in a single workflow. By securely connecting this multi-modal intelligence to your enterprise data wherever it lives, your engineering teams can turn this information into contextual business experiences.&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-video"&gt;



&lt;div class="article-module article-video "&gt;
  &lt;figure&gt;
    &lt;a class="h-c-video h-c-video--marquee"
      href="https://youtube.com/watch?v=kiuXDNR_d34"
      data-glue-modal-trigger="uni-modal-kiuXDNR_d34-"
      data-glue-modal-disabled-on-mobile="true"&gt;

      
        

        &lt;div class="article-video__aspect-image"
          style="background-image: url(https://storage.googleapis.com/gweb-cloudblog-publish/images/image2_NKCS8mx.max-1000x1000.png);"&gt;
          &lt;span class="h-u-visually-hidden"&gt;New Way Now: Kohl’s turns data analysis into active conversation with Gemini Enterprise&lt;/span&gt;
        &lt;/div&gt;
      
      &lt;svg role="img" class="h-c-video__play h-c-icon h-c-icon--color-white"&gt;
        &lt;use xlink:href="#mi-youtube-icon"&gt;&lt;/use&gt;
      &lt;/svg&gt;
    &lt;/a&gt;

    
      &lt;figcaption class="article-video__caption h-c-page"&gt;
        
          &lt;h4 class="h-c-headline h-c-headline--four h-u-font-weight-medium h-u-mt-std"&gt;See how Kohl’s uses Gemini Enterprise to turn data analysis into active conversations and deliver personalized customer experiences&lt;/h4&gt;
        
        
      &lt;/figcaption&gt;
    
  &lt;/figure&gt;
&lt;/div&gt;

&lt;div class="h-c-modal--video"
     data-glue-modal="uni-modal-kiuXDNR_d34-"
     data-glue-modal-close-label="Close Dialog"&gt;
   &lt;a class="glue-yt-video"
      data-glue-yt-video-autoplay="true"
      data-glue-yt-video-height="99%"
      data-glue-yt-video-vid="kiuXDNR_d34"
      data-glue-yt-video-width="100%"
      href="https://youtube.com/watch?v=kiuXDNR_d34"
      ng-cloak&gt;
   &lt;/a&gt;
&lt;/div&gt;

&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;h3&gt;&lt;strong style="vertical-align: baseline;"&gt;Enterprise trust, built in&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Trust is foundational to AI adoption, which is why Agent Platform is built with enterprise-grade governance, security, and observability by default. To help you scale with confidence, Agent Platform features built-in guardrails, continuous evaluation, and real-time tracking for costs, latency, and token usage. This operational transparency gives organizations the control they need to safely expand agentic workflows across the business. Builders can also leverage capabilities like &lt;/span&gt;&lt;a href="https://cloud.google.com/products/knowledge-catalog"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Knowledge Catalog&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; to establish a universal context engine across the enterprise, aggregating metadata with zero-copy federation to improve agent accuracy. &lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-video"&gt;



&lt;div class="article-module article-video "&gt;
  &lt;figure&gt;
    &lt;a class="h-c-video h-c-video--marquee"
      href="https://youtube.com/watch?v=Xynv0oBjvsw"
      data-glue-modal-trigger="uni-modal-Xynv0oBjvsw-"
      data-glue-modal-disabled-on-mobile="true"&gt;

      
        

        &lt;div class="article-video__aspect-image"
          style="background-image: url(https://storage.googleapis.com/gweb-cloudblog-publish/images/image3_zeEzYZA.max-1000x1000.png);"&gt;
          &lt;span class="h-u-visually-hidden"&gt;New Way Now: TELUS builds an Agentic Data Cloud to turn connectivity into customer intelligence&lt;/span&gt;
        &lt;/div&gt;
      
      &lt;svg role="img" class="h-c-video__play h-c-icon h-c-icon--color-white"&gt;
        &lt;use xlink:href="#mi-youtube-icon"&gt;&lt;/use&gt;
      &lt;/svg&gt;
    &lt;/a&gt;

    
      &lt;figcaption class="article-video__caption h-c-page"&gt;
        
          &lt;h4 class="h-c-headline h-c-headline--four h-u-font-weight-medium h-u-mt-std"&gt;See how TELUS uses Gemini Enterprise to turn data silos into an Agentic Data Cloud, driving real-time, proactive customer experiences.&lt;/h4&gt;
        
        
      &lt;/figcaption&gt;
    
  &lt;/figure&gt;
&lt;/div&gt;

&lt;div class="h-c-modal--video"
     data-glue-modal="uni-modal-Xynv0oBjvsw-"
     data-glue-modal-close-label="Close Dialog"&gt;
   &lt;a class="glue-yt-video"
      data-glue-yt-video-autoplay="true"
      data-glue-yt-video-height="99%"
      data-glue-yt-video-vid="Xynv0oBjvsw"
      data-glue-yt-video-width="100%"
      href="https://youtube.com/watch?v=Xynv0oBjvsw"
      ng-cloak&gt;
   &lt;/a&gt;
&lt;/div&gt;

&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;Looking ahead&lt;/span&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;We believe this recognition reflects our commitment to delivering an open, scalable, and powerful AI platform. As agentic systems reshape how enterprise software is built, Gemini Enterprise will continue to provide the foundation developers need to build with confidence.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Want to dive deeper into the report? Access the full Forrester Wave™: AI Platforms, Q3 2026 report &lt;/span&gt;&lt;a href="https://cloud.google.com/resources/content/forrester-wave-aiplatforms-2026"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;here&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;.&lt;/span&gt;&lt;/p&gt;
&lt;hr/&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Forrester does not endorse any company, product, brand, or service included in its research publications and does not advise any person to select the products or services of any company or brand based on the ratings included in such publications. Information is based on the best available resources. Opinions reflect judgment at the time and are subject to change. This report is part of a broader collection of Forrester resources, including interactive models, frameworks, tools, data, and access to analyst guidance. For more information, read about Forrester’s objectivity &lt;/span&gt;&lt;a href="https://www.forrester.com/about-us/objectivity/" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;here&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;.&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;</description><pubDate>Mon, 10 Aug 2026 16:00:00 +0000</pubDate><guid>https://cloud.google.com/blog/products/ai-machine-learning/google-named-a-leader-in-the-forrester-wave-ai-platforms/</guid><category>AI &amp; Machine Learning</category><og xmlns:og="http://ogp.me/ns#"><type>article</type><title>Google named a Leader in The Forrester Wave™: AI Platforms, Q3 2026</title><description></description><site_name>Google</site_name><url>https://cloud.google.com/blog/products/ai-machine-learning/google-named-a-leader-in-the-forrester-wave-ai-platforms/</url></og><author xmlns:author="http://www.w3.org/2005/Atom"><name>Jamie de Guerre</name><title>Senior Director, Outbound Product Management</title><department></department><company></company></author></item><item><title>Your agentic summer: No-cost lessons from Google experts to build and scale agents</title><link>https://cloud.google.com/blog/topics/training-certifications/free-gemini-enterrprise-training/</link><description>&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;I’ve talked to developers, IT leaders, and builders who all ask the same question: How do we actually get agents into production? &lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;The answer isn't theoretical — it's hands-on. Whether it’s designing a system that allows your agents to interact with external data sources while maintaining strict security guardrails or creating self-optimizing supply chain workflows or whatever you can think up, we’ve got you covered.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;That’s why we’ve designed a path to help you take your AI ideas from a rough sketch to fully autonomous agents running in production. This summer, you can harness the same frameworks and approaches used by Google experts to build and scale agents — &lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;entirely at no cost. &lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Powered by &lt;/span&gt;&lt;a href="https://developers.google.com/program/gear" rel="noopener" target="_blank"&gt;&lt;strong style="text-decoration: underline; vertical-align: baseline;"&gt;Gemini Enterprise Agent Ready (GEAR)&lt;/strong&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;, these hands-on labs and courses give you the blueprints and tools you need to deploy agents that ship&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt; Find your roadmap to future-proof your skills this summer, starting here.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;1. &lt;/strong&gt;&lt;a href="https://www.skills.google/paths/3546/course_templates/1583" rel="noopener" target="_blank"&gt;&lt;strong style="text-decoration: underline; vertical-align: baseline;"&gt;Intro to AI Agents&lt;/strong&gt;&lt;/a&gt;&lt;strong style="vertical-align: baseline;"&gt;: &lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;Build a foundational understanding of how autonomous agents can redefine productivity. &lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;2. &lt;/strong&gt;&lt;a href="https://www.skills.google/paths/3546/course_templates/1562" rel="noopener" target="_blank"&gt;&lt;strong style="text-decoration: underline; vertical-align: baseline;"&gt;Agent Fundamentals&lt;/strong&gt;&lt;/a&gt;&lt;strong style="vertical-align: baseline;"&gt;: &lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;Go under the hood of autonomous intelligence. Learn decision models and execution loops to deploy adaptive agents over rigid automation. &lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;3. &lt;/strong&gt;&lt;a href="https://www.skills.google/paths/3546/course_templates/1587" rel="noopener" target="_blank"&gt;&lt;strong style="text-decoration: underline; vertical-align: baseline;"&gt;Enterprise Agents and Use Cases&lt;/strong&gt;&lt;/a&gt;&lt;strong style="vertical-align: baseline;"&gt;: &lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;Discover how AI agents drive real business impact. Map agents directly to corporate KPIs, solve operational bottlenecks, and utilize no-code to high-code frameworks.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;4. &lt;/strong&gt;&lt;a href="https://www.skills.google/paths/3546/course_templates/1586" rel="noopener" target="_blank"&gt;&lt;strong style="text-decoration: underline; vertical-align: baseline;"&gt;Create Your First Gemini Enterprise Application skill badge&lt;/strong&gt;&lt;/a&gt;&lt;strong style="vertical-align: baseline;"&gt;: &lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;Earn a skill badge that proves you can create an app with Gemini Enterprise. You will master &lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;capabilities like deep research agents, multi-agent ideation, and Gemini Notebook for focused analysis.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;5. &lt;/strong&gt;&lt;a href="https://www.skills.google/paths/3980/course_templates/1643" rel="noopener" target="_blank"&gt;&lt;strong style="text-decoration: underline; vertical-align: baseline;"&gt;Human-Centered AI&lt;/strong&gt;&lt;/a&gt;&lt;strong style="vertical-align: baseline;"&gt;: &lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;Keep humanity at the core of automation. Learn to strategically balance machine speed with human intuition for successful orchestration. &lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;6. &lt;/strong&gt;&lt;a href="https://www.skills.google/course_templates/1672" rel="noopener" target="_blank"&gt;&lt;strong style="text-decoration: underline; vertical-align: baseline;"&gt;Agentic Strategy: Discover, Design, and Prototype&lt;/strong&gt;&lt;/a&gt;&lt;strong style="vertical-align: baseline;"&gt;: &lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;Prototype high-impact AI projects with zero code. Leverage Google’s transformation framework, map user journeys and build functional retail prototypes.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;7. &lt;/strong&gt;&lt;a href="https://www.skills.google/paths/3980/course_templates/1682" rel="noopener" target="_blank"&gt;&lt;strong style="text-decoration: underline; vertical-align: baseline;"&gt;Orchestrate Multi-Agent Workflows with Gemini Enterprise skill badge&lt;/strong&gt;&lt;/a&gt;&lt;strong style="vertical-align: baseline;"&gt;: &lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;Demonstrate your ability to manage multiple agents powered by Gemini Enterprise with a skill badge. This skill badge shows that you can unify data across first- and third-party sources, develop multimedia marketing materials, and fully automate complex business actions across disjointed systems.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;8. &lt;/strong&gt;&lt;a href="https://www.skills.google/paths/3545/course_templates/1596/labs/618257" rel="noopener" target="_blank"&gt;&lt;strong style="text-decoration: underline; vertical-align: baseline;"&gt;Engineer AI Agents with Agent Development Kit (ADK) skill badge:&lt;/strong&gt;&lt;/a&gt;&lt;strong style="vertical-align: baseline;"&gt; &lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;Build production-grade agents using expert developer tools. Earn a skill badge that proves you can perform live search grounding, build structured JSON schemas, and manage ADK pipelines.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;9. &lt;/strong&gt;&lt;a href="https://www.skills.google/focuses/143124?parent=catalog&amp;amp;path=3545" rel="noopener" target="_blank"&gt;&lt;strong style="text-decoration: underline; vertical-align: baseline;"&gt;Add Currency Tools to an Agent Using MCP&lt;/strong&gt;&lt;/a&gt;&lt;strong style="vertical-align: baseline;"&gt;:&lt;/strong&gt;&lt;strong style="vertical-align: baseline;"&gt; &lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;Connect your LLMs to external systems in just 20 minutes. Securely bridge agents with live external databases and deploy via CLI.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;10. &lt;/strong&gt;&lt;a href="https://www.skills.google/paths/3545/course_templates/1584" rel="noopener" target="_blank"&gt;&lt;strong style="text-decoration: underline; vertical-align: baseline;"&gt;Manage Agent Memory and State&lt;/strong&gt;&lt;/a&gt;&lt;strong style="vertical-align: baseline;"&gt;: &lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;Give your agents a memory. Move beyond single-query replies and use session states with the ADK to build highly personalized, deeply contextual agents.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;11. &lt;/strong&gt;&lt;a href="https://www.skills.google/course_templates/1832" rel="noopener" target="_blank"&gt;&lt;strong style="text-decoration: underline; vertical-align: baseline;"&gt;Create Agent Skills with Google&lt;/strong&gt;&lt;/a&gt;&lt;strong style="vertical-align: baseline;"&gt;: &lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;Infuse domain expertise into custom skills. Minimize AI unpredictability and build reusable workflows that optimize agent performance.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;12. &lt;/strong&gt;&lt;a href="https://www.skills.google/course_templates/1782?catalog_rank=%7B%22rank%22%3A1%2C%22num_filters%22%3A0%2C%22has_search%22%3Atrue%7D&amp;amp;search_id=91630203" rel="noopener" target="_blank"&gt;&lt;strong style="text-decoration: underline; vertical-align: baseline;"&gt;AgentOps: Operationalize AI Agents on Google Cloud&lt;/strong&gt;&lt;/a&gt;&lt;strong style="vertical-align: baseline;"&gt;: &lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;Harden your prototypes and scale safely to production. Implement observability, proactive monitoring dashboards, and robust CI/CD security.&lt;/span&gt;&lt;/p&gt;
&lt;h3&gt;&lt;strong style="vertical-align: baseline;"&gt;Test your skills at the summertime Hackathon&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Keep moving with agents! The &lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;All Things Agentic&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Hackathon is officially live.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;The next leap in AI won't build itself — it needs you. Step up to the challenge with Gemini 3.5 and Google Cloud and deploy autonomous agents that do the heavy lifting in the background. Build what’s next, show the world what you can do, and compete for $180,000 in prizes, cash, and credits&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Submissions are open from August 3, 2026 to August 31, 2026. Register &lt;/span&gt;&lt;a href="http://allthingsagentichackathon.devpost.com" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;here.&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; &lt;/span&gt;&lt;/p&gt;
&lt;h3&gt;&lt;strong style="vertical-align: baseline;"&gt;Join GEAR today&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Don't wait for the summer to pass you by. Get hands-on with the tools, earn real-world credentials, and build in-demand skills, with confidence.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Ready to level up your agentic skills? Check out our two newest learning paths: &lt;/span&gt;&lt;a href="https://www.skills.google/paths/4459" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Build High-Performance Multi-Agent Systems&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; and &lt;/span&gt;&lt;a href="https://www.skills.google/paths/4461" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Govern and Secure Enterprise Agents&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;. &lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Learn more → &lt;/span&gt;&lt;a href="https://developers.google.com/program/gear" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Join GEAR today&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; and start building.&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;</description><pubDate>Thu, 06 Aug 2026 16:00:00 +0000</pubDate><guid>https://cloud.google.com/blog/topics/training-certifications/free-gemini-enterrprise-training/</guid><category>AI &amp; Machine Learning</category><category>Training and Certifications</category><media:content height="540" url="https://storage.googleapis.com/gweb-cloudblog-publish/images/linkedin_header.max-600x600.png" width="540"></media:content><og xmlns:og="http://ogp.me/ns#"><type>article</type><title>Your agentic summer: No-cost lessons from Google experts to build and scale agents</title><description></description><image>https://storage.googleapis.com/gweb-cloudblog-publish/images/linkedin_header.max-600x600.png</image><site_name>Google</site_name><url>https://cloud.google.com/blog/topics/training-certifications/free-gemini-enterrprise-training/</url></og><author xmlns:author="http://www.w3.org/2005/Atom"><name>Gary Eimerman</name><title>Managing Director, Google Cloud Learning</title><department></department><company></company></author></item><item><title>Mirendil taps AI Hypercomputer TPUs and GPUs for pre- and post-training applications</title><link>https://cloud.google.com/blog/topics/startups/mirendil-selects-ai-hypercomputer/</link><description>&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Nearly every major AI lab uses Google Cloud infrastructure, including for training of models, inference for agents, and new frontier research. Google Cloud also continues to be the platform of choice for new, high-growth AI startups who are driving much of the industry’s research and innovation.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Today, we’re announcing that &lt;/span&gt;&lt;a href="https://mirendil.com/" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Mirendil&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;, an exciting frontier AI lab focused on accelerating AI development, will also utilize Google Cloud’s &lt;/span&gt;&lt;a href="https://cloud.google.com/ai-infrastructure"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;AI Hypercomputer&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;. This includes using a mix of Google’s TPU AI accelerators and full-stack NVIDIA AI infrastructure running on Google Cloud; this purpose-built AI infrastructure will support model pre-training and post-training applications for Mirendil. &lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;The Mirendil team is building new AI systems that can help accelerate and democratize AI research and development. This means managing complex, end-to-end training workflows from initial model pre-training through post-training, and powering reinforcement learning on a massive scale. The ability to choose a mix of both TPU and NVIDIA’s full-stack accelerated computing platform through Google Cloud meant that Mirendil could access critical compute very quickly, and continue to match its workloads to the architecture best-suited to it over time.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;We closely partnered with Mirendil on end-to-end design and deployment of combined TPU and NVIDIA AI infrastructure across compute, storage, networking, and control planes. We also collaborated on a system that uses &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/gemini-enterprise-agent-platform/machine-learning/training/training-clusters/overview"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;managed training clusters running in Gemini Enterprise Agent Platform&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;, which effectively streamlines the provisioning and management of both TPU and GPU environments for Mirendil. Mirendil is already live with a cluster of TPU v5P chips, with NVIDIA AI accelerated computing systems coming online soon.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;&lt;span style="vertical-align: baseline;"&gt;"Progress in AI has been bounded by how fast humans can run the research loop - designing experiments, evaluating results, and iterating," said Behnam Neyshabur, cofounder and CEO of Mirendil. "We're building AI systems that can accelerate and improve that loop itself. Expanding on Google Cloud gives us the scale and flexibility to push those systems further and put frontier AI research capabilities in the hands of many more scientists and engineers to run that loop faster and at a greater scale."&lt;/span&gt;&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;&lt;span style="vertical-align: baseline;"&gt;&lt;span style="vertical-align: baseline;"&gt;You can read more about our partnership on Mirendil’s &lt;/span&gt;&lt;a href="https://mirendil.com/news/scaling-self-accelerating-ai-with-google/" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;blog&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;.&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;</description><pubDate>Thu, 06 Aug 2026 13:00:00 +0000</pubDate><guid>https://cloud.google.com/blog/topics/startups/mirendil-selects-ai-hypercomputer/</guid><category>AI &amp; Machine Learning</category><category>AI infrastructure</category><category>Customers</category><category>Startups</category><media:content height="540" url="https://storage.googleapis.com/gweb-cloudblog-publish/images/mirendil.max-600x600.jpg" width="540"></media:content><og xmlns:og="http://ogp.me/ns#"><type>article</type><title>Mirendil taps AI Hypercomputer TPUs and GPUs for pre- and post-training applications</title><description></description><image>https://storage.googleapis.com/gweb-cloudblog-publish/images/mirendil.max-600x600.jpg</image><site_name>Google</site_name><url>https://cloud.google.com/blog/topics/startups/mirendil-selects-ai-hypercomputer/</url></og><author xmlns:author="http://www.w3.org/2005/Atom"><name>Darren Mowry</name><title>VP, Global Startups and Investor Ecosystem, Google</title><department></department><company></company></author></item></channel></rss>