惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

云风的 BLOG
云风的 BLOG
博客园 - 三生石上(FineUI控件)
WordPress大学
WordPress大学
F
Fortinet All Blogs
奇客Solidot–传递最新科技情报
奇客Solidot–传递最新科技情报
博客园 - 叶小钗
爱范儿
爱范儿
美团技术团队
H
Hackread – Cybersecurity News, Data Breaches, AI and More
有赞技术团队
有赞技术团队
博客园_首页
T
The Blog of Author Tim Ferriss
T
Tailwind CSS Blog
V
Visual Studio Blog
Jina AI
Jina AI
博客园 - Franky
量子位
MongoDB | Blog
MongoDB | Blog
L
LangChain Blog
Apple Machine Learning Research
Apple Machine Learning Research
freeCodeCamp Programming Tutorials: Python, JavaScript, Git & More
U
Unit 42
aimingoo的专栏
aimingoo的专栏
M
MIT News - Artificial intelligence

Google Developers Blog

Build zero-trust AI agents that judge intent, not just syntax- Google Developers Blog Autonomous LLM post-training with Tunix on TPUs- Google Developers Blog The Anatomy of Harness Engineering: How to Evaluate, Iterate, and Guard AI Coding Agents- Google Developers Blog Announcing ADK for Kotlin 1.0: Building Production-Ready AI Agents in Kotlin, Android, and Beyond- Google Developers Blog Driving Developer Excellence: Inside the Program Sprints- Google Developers Blog 4 engineering patterns behind the strongest AI Agents Challenge submissions- Google Developers Blog Decoding cosmic signals with deep learning and Keras- Google Developers Blog How to Evaluate Live & Voice Agents in ADK- Google Developers Blog Build zero-trust AI agents with Google's Agent Development Kit- Google Developers Blog Introducing Credentio: Open Source C++ Library for C2PA Content Credentials from Google- Google Developers Blog HeyGen x Google Cloud: Bringing Avatar IV to TPUs- Google Developers Blog Why Go is an Ideal Language for AI-Assisted Software Engineering- Google Developers Blog Mastering Edge AI on Raspberry Pi with LiteRT and Gemma- Google Developers Blog Agent Plugins package your skills, tools, and more- Google Developers Blog Scaling AI Agent Infrastructure with the MCP Stateless updates- Google Developers Blog A unified API for AI model routing- Google Developers Blog Scaling real-time AI agents with session-aware load balancing- Google Developers Blog Agent and Model Evaluations in Gemini Enterprise Agent Platform are now GA- Google Developers Blog Enable on-demand expertise with Agent Skills in Genkit Go- Google Developers Blog How to use Google microbenchmarks for evaluating TPU performance- Google Developers Blog Run Ray on TPU, Part 2: Ray AI libraries- Google Developers Blog Scaling Agentic RL: High-Throughput Agentic Training with Tunix- Google Developers Blog Run Ray on TPU, Part 1: The foundations- Google Developers Blog Expanding Choice in Gemini Enterprise Agent Platform: Introducing Grounding with Parallel Web Search- Google Developers Blog Building scalable AI agents with modular prompt transpilation- Google Developers Blog Evolving Spec-Driven Development: Conductor Now Supports Antigravity- Google Developers Blog Systems Engineering Playbook: Optimizing Qwen 3.5-397B MoE on Ironwood (TPU7x)- Google Developers Blog Unlocking the Next Era of On-Device AI with Google Tensor and Pixel- Google Developers Blog LiteRT.js, Google's high performance Web AI Inference- Google Developers Blog Bridging the Domain Gap: AI Race Coach built with Antigravity and Gemini- Google Developers Blog
Enterprise-Grade Precision for Long-Context Multimodal Em...
Anthony Su, Injae Kwak · 2026-08-26 · via Google Developers Blog

What is an Embedding Model and What is it Used For?

In modern AI architectures, embedding models serve as the foundational translators bridging raw unstructured data and downstream intelligent reasoning. Simply put, embedding models translate inputs of various types of data — including text, images, and audio — into dense vector math. These high-dimensional numeric arrays capture semantic relationships, powering critical enterprise capabilities such as semantic search, recommender systems, intent classification, personalized content discovery, and vector clustering.

image5

Figure 1. Embedding models enterprise capabilities.

To understand how embedding models work in practice, consider a standard vector search query. When a user queries a semantic search engine for the word "cat", the model maps the token into a dense coordinate space. In this vector space, the mathematical distance between "cat" and "feline" or "dog" is short, yielding high similarity scores; conversely, terms like "hat" or "car" map to distant coordinates despite their orthographic similarity.

Seamless Elasticity with vLLM-TPU & GKE

While deploying small text embedding models for prototype applications is straightforward, scaling pipelines to serve millions of queries introduces different types of production bottlenecks. The most common ones we see are accessing elastic capacity of accelerators to seamlessly scale compute resources alongside dynamic traffic fluctuations and improving cost/performance efficiency.

To overcome these scaling and capacity constraints, Google Cloud has integrated native TPU support into vLLM — the industry-standard, highly optimized and popular open-source LLM serving engine. Standardizing on vLLM for TPU serving provides architecture true elasticity. Engineering teams can scale serving capacity up and down dynamically by provisioning TPU nodes directly alongside other XPU instances.

image6 (1)

Figure 2. If primary TPU reservations are fully utilized, the serving infrastructure automatically falls back to secondary GPU spot or on-demand pools without interrupting incoming inference traffic.

By taking advantage of primitives like Custom Compute Classes in Google Kubernetes Engine (GKE), organizations can automate node autoscaling based on strict priority rules to scale up across different capacity types or accelerators if the previous one isn’t available.

Engineering High-Precision Embedding Support on TPUs

Serving next-generation embedding models in production demands processing ultra-long sequence contexts - ranging from 4K+ tokens for text workloads up to 15K+ tokens for multimodal text-and-image inputs. Crucially, enterprise applications require that these embeddings maintain strict mathematical parity and high precision across heterogeneous hardware backends compared to reference.

To bring high-dimensional vector pooling models to TPU hardware topologies, we took the Qwen3 Embedding model series as the target engineering models and engineered several key optimizations of the vLLM framework on TPU.

Challenge A: Hardware-Safe Tensor Alignment

TPU Matrix Execution Units (MXUs) impose strict divisibility constraints when sharding vocabulary matrices across topology meshes via Tensor Parallelism (TP). We implemented a unified, hardware-safe vocabulary padding strategy that guarantees exact tensor alignment during All-Gather execution.

Challenge B: Materialization Hardening, TPU Lazy-Loading & Compilation Pre-warming

vLLM relies on lazy-loading mechanisms on TPUs to minimize server cold-start latencies and reduce host memory peaks. To eliminate model initialization failures during lazy tensor transformations, we introduced attribute promotion within the unquantization pipeline, making weight loading fully compatible with vLLM’s TPU lazy-loader for zero-failure initialization.

Furthermore, to eliminate runtime JIT compilation latencies and avoid compilation traps in multi-processes deployments, we implemented sharding-aware pre-warming to lock JAX/XLA compilation caches prior to inference and stabilize the production pipelines and rollouts.

Challenge C: Long-Context StepPool Architecture

Ultra-long contexts require Chunked Prefill in the pooling layer to prevent High Bandwidth Memory (HBM) exhaustion, creating risk of state loss across step boundaries. We engineered a hybrid StepPool and migrated metadata to CachedRequestState, ensuring pooling states correctly accumulate across steps and survive request preemptions.

vLLM Embedding Sample on TPU

Below is a minimal example demonstrating how to initialize Qwen3-Embedding-8B on TPU. For complete setup scripts and environment deployment steps, refer to the official AI-Hypercomputer Qwen3-Embedding-8B Recipes on GitHub:

from vllm import LLM

# Initialize Qwen3-Embedding-8B on Cloud TPU using vLLM's native pooling runner
llm = LLM(
    model="Qwen/Qwen3-Embedding-8B",
    runner="pooling",             # Enables dense pooling output
    tensor_parallel_size=2,       # Sharded across TPU topology mesh
    max_model_len=16384,
    max_num_batched_tokens=512,
    dtype="bfloat16",
    trust_remote_code=True
)

# Extract dense vector embeddings across inputs
prompts = ["Enterprise-grade semantic retrieval on TPUs with vLLM."]
results = llm.embed(prompts)
embedding_vector = results[0].outputs.embedding

Python

Copied

Golden-Reference Precision & Numerical Parity

To certify enterprise-grade precision, we conducted rigorous mathematical parity evaluations comparing TPU outputs against other XPU golden references across multi-language and multimodal datasets.

To evaluate numerical alignment between dense embedding vectors generated on TPUs (vTpu) and reference baseline vectors generated on XPUs (vRef), we calculate their cosine similarity:

image3 (1)

A cosine similarity score approaching 1.0 (with a target quality pass threshold of ≥0.999 for text and ≥0.995 for multimodal inputs) demonstrates near-perfect numerical parity across hardware backends. This confirms that optimizations implemented on the vLLM-TPU stack maintain golden-reference precision without sacrificing accuracy.

For step-by-step instructions on generating pairwise calculations, check out the official AI-Hypercomputer Qwen3-Embedding-8B Recipes on GitHub.

Qwen3-Embedding-8B (7K+ Tokens, TPU vs. CPU Baseline)

table 1

Table1. While maintaining strict numerical alignment, serving qwen-3-embedding-8b (bf16, 16K+ sequence length, TP=4) on TPU Ironwood achieved an impressive throughput of 83,996 total token/s and 5.13 req/s.

Qwen3-VL-Embedding-8B (15K+ Tokens, TPU vs. XPU Baseline)

table 2

Table2. vLLM-TPU only chunks the text portion of multimodal prefill.

Public Recipes & Explore Resources

To help developers reproduce our numerical parity evaluations and rapidly deploy embedding workloads on Google Cloud TPUs, we have open-sourced official setup and execution recipes on the AI-Hypercomputer Public Repository.

Get Started:

* Qwen3-Embedding-8B TPU Recipe: Explore text embedding recipes at AI-Hypercomputer/Qwen3-Embedding-8B

* Qwen3-VL-Embedding-8B TPU Recipe: Explore multimodal embedding recipes at AI-Hypercomputer/Qwen3-VL-Embedding-8B

* vLLM TPU Engine: Explore more on vLLM framework on TPU at vllm-project/tpu-inference

* Google Cloud TPU Portal: Provision Cloud TPU instances and explore hardware specs at cloud.google.com/tpu

Acknowledgments

The engineering achievements and cross-hardware optimizations highlighted in this post were made possible through the incredible collaboration across Google Cloud Product, Engineering and the vLLM community.