惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

Hacker News: Ask HN
Hacker News: Ask HN
Recent Commits to openclaw:main
Recent Commits to openclaw:main
cs.CV updates on arXiv.org
cs.CV updates on arXiv.org
C
Check Point Blog
S
Security Affairs
Hacker News - Newest:
Hacker News - Newest: "LLM"
S
Secure Thoughts
Recorded Future
Recorded Future
cs.AI updates on arXiv.org
cs.AI updates on arXiv.org
T
The Blog of Author Tim Ferriss
B
Blog
C
Cybersecurity and Infrastructure Security Agency CISA
Google DeepMind News
Google DeepMind News
CTFtime.org: upcoming CTF events
CTFtime.org: upcoming CTF events
A
Arctic Wolf
T
The Exploit Database - CXSecurity.com
Stack Overflow Blog
Stack Overflow Blog
T
Threat Research - Cisco Blogs
GbyAI
GbyAI
AWS News Blog
AWS News Blog
MongoDB | Blog
MongoDB | Blog
Y
Y Combinator Blog
Google Online Security Blog
Google Online Security Blog
T
Troy Hunt's Blog
I
InfoQ
L
LINUX DO - 热门话题
WordPress大学
WordPress大学
C
Cisco Blogs
G
GRAHAM CLULEY
The Register - Security
The Register - Security
A
About on SuperTechFans
让小产品的独立变现更简单 - ezindie.com
让小产品的独立变现更简单 - ezindie.com
Schneier on Security
Schneier on Security
Project Zero
Project Zero
H
Hackread – Cybersecurity News, Data Breaches, AI and More
P
Privacy & Cybersecurity Law Blog
Cloudbric
Cloudbric
H
Hacker News: Front Page
小众软件
小众软件
雷峰网
雷峰网
The Hacker News
The Hacker News
www.infosecurity-magazine.com
www.infosecurity-magazine.com
T
Tor Project blog
博客园 - 聂微东
N
Netflix TechBlog - Medium
V
Vulnerabilities – Threatpost
The GitHub Blog
The GitHub Blog
腾讯CDC
P
Palo Alto Networks Blog
Scott Helme
Scott Helme

Crazyrouter Blog (English)

Ideogram AI Guide 2026: Product Mockups, Text Rendering, and API Automation Akool AI Voice Generator Review 2026: API Alternatives for Developers GLM 4.6 API Guide 2026: Build Chinese-English Agents with Tool Calling Google Veo3 API Guide 2026: Batch Video Generation, QA, and Fallbacks AI Lip Sync Tools Comparison 2026: Developer Guide for Localization Pipelines Claude Opus 4.8 vs Opus 4.7: Real API Benchmark Results for Developers Opus 4.8 vs Opus 4.7 Coding Test: What Changed for Developers? Opus 4.8 vs Opus 4.7 for Agents: JSON, Tool Use, and Structured Output Gemini 2.5 Flash-Lite for RAG, Agent Routing, and Cost per Successful Task Gemini 2.5 Flash-Lite for Support Automation and Ticket Triage Gemini 2.5 Flash-Lite Use Cases: The Practical Automation Tier for Developers Claude Jupiter v1-p vs GPT-5.5 Benchmark: Real API Test on Reasoning and Coding Claude Jupiter v1-p vs Claude Opus 4.7 vs Sonnet 4.6: Live API Test Claude Jupiter v1-p vs Claude Opus 4.7 vs Sonnet 4.6: Live API Test Claude Code Pricing 2026: Pro vs Max vs Team vs API Costs Claude Opus 4.7 vs DeepSeek V4 Pro: Real API Compatibility and Coding Benchmark Gemini CLI Complete Guide 2026: Repo Automation, CI Agents, and Multi-Model Routing Ideogram AI Guide 2026: Brand Design Automation, API Workflows, and Alternatives GLM 4.6 API Guide 2026: Agents, RAG, Tool Calling, and Bilingual Apps WAN 2.2 Animate Tutorial 2026: Character Consistency, Shot Control, and API Workflows Google Veo3 API Guide 2026: Production Video Pipelines, Prompts, Pricing, and Fallbacks AI API Pricing Comparison 2026: Text, Image, Video, Caching, and Router Costs Codex CLI Installation Guide 2026: Windows, macOS, Linux, Proxies, and CI Setup How to Get a Claude API Key in 2026: Secure Setup for Teams, CI, and Alternatives Gemini Advanced Review 2026: Is It Worth It for Coding, Research, and API Teams? Claude Code Pricing Guide 2026: Team Agent Budgets, API Fallbacks, and Cost Control Seedance 2.0 Pricing: Convert 46 CNY per Million Tokens to Cost per Second Qwen2.5-Omni Guide 2026: Real-Time Voice, Vision, and Multimodal Agents Kimi K2 Thinking Guide 2026: Reasoning Workflows, Evals, and Cost Control Google Veo3 API Guide 2026: Batch Video Pipelines, Pricing, and Fallbacks Codex CLI Installation Guide 2026: macOS, Linux, WSL, Proxies, and Dev Containers How to Get a Claude API Key in 2026: Safe Production Setup and Alternatives AI API Pricing Comparison 2026: GPT, Claude, Gemini, Video, and Agent Workloads Gemini Advanced Review 2026: Is It Worth It for Developer Teams? Claude Code Pricing Guide 2026: API Fallbacks, Team Seats, and Budget Control Seedream 4.0 API Tutorial 2026: Batch Image Generation, Product Creative, and Pricing Qwen2.5-Omni Guide 2026: Real-Time Voice, Vision, Text Agents, and API Integration Kimi K2 Thinking Guide 2026: Reasoning Agents, Evaluation Workflows, and API Cost Control WAN 2.2 Animate Tutorial 2026: Character Motion, Shot Control, API Pipelines, and Pricing Google Veo3 API Guide 2026: Production Video Workflows, Prompts, Pricing, and Fallbacks AI API Pricing Comparison 2026: OpenAI, Claude, Gemini, DeepSeek, and Router Costs How to Get a Claude API Key in 2026: Setup, Security, Rotation, and Alternatives Codex CLI Installation Guide 2026: macOS, Linux, WSL, Proxies, and Devcontainers Gemini Advanced Review 2026: Is It Worth It for Developers and API Builders? Claude Code Pricing Guide 2026: CI Agents, Team Seats, and API Budget Planning AI API Gateway for Singapore and Malaysia Developers: One Endpoint for GPT, Claude and Gemini AI API Gateway for Thai Developers: Use GPT, Claude and Gemini with One Key One API Key for GPT, Claude and Gemini: A Practical Setup for Central Asia Developers Gemini 3.5 Flash vs Claude Response-Tier Models: Which One Should Developers Use? Gemini 3.5 Flash vs Gemini 3 Flash vs Gemini 2.5 Flash: Real API Benchmark "How to Test Multiple AI Image Models with One API Key" Codex CLI Installation Guide: Setup on macOS, Linux, Windows WSL and CI/CD Seedream 4.0 API Tutorial: ByteDance Image Generation for Production Pipelines Kimi K2 Thinking Model: Complete Developer Guide for Reasoning Workflows Luma Ray 2 Review: AI Video Generation Quality, Speed, and API Guide Pika 2.2 New Features Review: Scene Director, Sound Design, and API Updates Google Veo 3 API Guide: Video Generation with Audio for Developers AI Lip Sync Tools Comparison 2026: Best APIs for Talking Avatars and Video Dubbing Gemini Advanced Review May 2026: Is It Worth $20/Month for AI Power Users? Claude Code Pricing in May 2026: Max Plan, Opus 4, and Real Cost Breakdown Hermes Agent + Crazyrouter: One-Click Setup for 627+ AI Models AI Meme Generator & Coloring Book Creator with GPT-image-2 — Fun Projects That Actually Make Money AI Future Baby Prediction with GPT-image-2 — See What Your Child Might Look Like Ghibli Style Photo Transformation with GPT-image-2 — Turn Any Photo Into Anime Art AI Action Figure Generator with GPT-image-2 — Turn Anyone Into a Boxed Toy AI Face Reading & Personal Color Analysis with GPT-image-2 — Two Viral Use Cases in One Guide AI Palm Reading with GPT-image-2 — Generate Professional Palmistry Analysis from a Single Photo Gemini 2.5 Flash-Lite Pricing Explained — The Cheapest Gemini Model for High-Volume Workloads Claude Sonnet 4.6 Pricing Explained — Caching, Tiers, and How to Save 45% with Crazyrouter Gemini Free vs Gemini Advanced: Pricing, Limits, Features, and Is It Worth Paying For? AI Context Window Comparison (2026): GPT, Claude, Gemini Token Limits by Model Claude Sonnet 4.5 Pricing Explained — Caching, Batch API, and How to Save 45% with Crazyrouter Claude Opus 4.7 Pricing Explained — New Tokenizer, Caching, and How to Save 45% with Crazyrouter Claude Opus 4.6 Pricing Explained — Caching, Tiers, and How to Save 45% with Crazyrouter Best AI Models for RAG Applications 2026: Embeddings, Retrieval, and Generation Seedance 2.0 vs Kling 2.1 vs Runway Gen 4 Turbo: Video AI API Comparison 2026 AI Video Generation API Pricing May 2026: Veo3 vs Kling vs Runway vs Sora How to Get Claude API Key in China 2026: Complete Setup Guide AI Coding Tools ROI Calculator: Claude Code vs Codex CLI vs Gemini CLI Cost Analysis 2026 AI API Pricing Comparison May 2026 - Complete Developer Guide Grok 4 API Pricing Complete Guide 2026 DeepSeek R2: The 32B Reasoning Model That Runs on a Single GPU — Complete Guide for Developers "GPT-5.1 Codex Max Pricing Explained — The Code-Specialized Model and How to Save with Crazyrouter" GPT-4o Pricing Explained — The Legacy Flagship That's Still Worth Using GLM-5 Pricing Explained — Zhipu AI's Flagship Model and How to Access via Crazyrouter Gemini 3 Flash Pricing Explained — Balanced Speed and Cost with Crazyrouter Savings "Gemini 3.1 Pro Pricing Explained — Context Tiers, Caching, and How to Save with Crazyrouter" GPT-5.5 Pricing Explained — OpenAI's Latest Flagship, Reasoning Tokens, and How to Save with Crazyrouter AI Model Pricing Guide 2026: What Every Model Costs on Crazyrouter (and How Much You Save) MiniMax M2 Pricing Explained — China's Competitive AI Model and How to Access via Crazyrouter Grok 4.1 Thinking Pricing Explained — Reasoning Tokens, Caching, and How to Save with Crazyrouter Grok 4.1 Pricing Explained — 2M Context, Caching, Tool Costs, and How to Save with Crazyrouter GPT-5 Pricing Explained — Reasoning Tokens, Caching, Batch API, and How to Save with Crazyrouter GPT-5-nano Pricing Explained — The Cheapest GPT Model for High-Throughput Workloads GPT-5-mini Pricing Explained — Ultra-Low Cost AI with Caching and Batch Discounts GPT-5.4 Pricing Explained — Cached Input, Context Tiers, Batch API, and How to Save with Crazyrouter GPT-5.2 Pricing Explained — Caching, Batch API, and How to Save with Crazyrouter OpenRouter vs Crazyrouter (2026): Pricing, Models, and Which API Gateway Fits Developers Better Suno v4 vs v5 vs v4.5: Which Version Sounds Better and Is Worth Using in 2026? How to Use Claude Code with Crazyrouter: Base URL Setup, Model Routing, and Cost Savings
Text-Embedding-3-Small: Complete Guide to OpenAI's Most Popular Embedding Model (2026)
Crazyrouter Team · 2026-05-03 · via Crazyrouter Blog (English)

text-embedding-3-small is OpenAI's cost-effective embedding model, released in January 2024. It converts text into 1536-dimensional vectors that capture semantic meaning — the foundation for semantic search, RAG pipelines, recommendation systems, and classification tasks.

This guide covers everything: pricing, token limits, dimensions, API usage, dimension reduction, performance benchmarks, and how it compares to text-embedding-3-large.

Text-Embedding-3-Small Quick Reference#

SpecValue
Model nametext-embedding-3-small
ProviderOpenAI
Default dimensions1536
Adjustable dimensions256 – 1536
Max input tokens8,191
Max batch size2,048 inputs per request
Pricing (OpenAI direct)$0.020 per 1M tokens
Pricing (Crazyrouter)$0.016 per 1M tokens
MTEB benchmark score62.3
MultilingualYes
Output formatfloat or base64
Release dateJanuary 25, 2024
StatusActive (not deprecated)

Text-Embedding-3-Small Pricing#

text-embedding-3-small costs 0.020per1milliontokens∗∗onOpenAIdirectly.Through[Crazyrouter](https://crazyrouter.com),thepricedropsto∗∗0.020 per 1 million tokens** on OpenAI directly. Through [Crazyrouter](https://crazyrouter.com), the price drops to **0.016 per 1M tokens — a 20% discount.

To put that in perspective:

Document VolumeApprox. TokensCost (OpenAI)Cost (Crazyrouter)
100 pages of text~75,000$0.0015$0.0012
10,000 pages~7.5M$0.15$0.12
1 million pages~750M$15.00$12.00
Wikipedia (English, full)~4.4B$88.00$70.40

Cost Comparison with Other Embedding Models#

ModelPrice / 1M TokensPrice via CrazyrouterDimensions
text-embedding-3-small$0.020$0.0161536
text-embedding-3-large$0.130$0.1003072
text-embedding-ada-002$0.1001536
Google text-embedding-005$0.00625$0.005768
Cohere embed-v4$0.1001024
Voyage voyage-3-large$0.1802048

text-embedding-3-small is 6.5x cheaper than text-embedding-3-large and 5x cheaper than the older ada-002 — while outperforming ada-002 on benchmarks.

For a deeper comparison of all embedding models, see our AI Embeddings Comparison 2026 Guide.

Text-Embedding-3-Small Dimensions#

The default output is a 1536-dimensional vector. But text-embedding-3-small supports dimension reduction via the dimensions parameter — you can request any value from 256 to 1536.

This is done using Matryoshka Representation Learning (MRL). The model is trained so that the first N dimensions of the vector carry the most important information. Truncating to fewer dimensions loses some nuance but keeps most of the semantic signal.

Dimension vs. Quality Tradeoff#

DimensionsMTEB ScoreStorage per VectorRelative Quality
1536 (default)62.36,144 bytes100%
1024~61.54,096 bytes~98.7%
768~60.83,072 bytes~97.6%
512~59.72,048 bytes~95.8%
256~57.81,024 bytes~92.8%

When to Reduce Dimensions#

  • 256 dimensions: Prototyping, low-resource environments, or when storage is the bottleneck
  • 512 dimensions: Good balance for mobile apps or edge deployments
  • 768 dimensions: Matches Google's embedding size — useful for migration
  • 1536 dimensions: Production workloads where quality matters most

Text-Embedding-3-Small Token Limit and Context Length#

text-embedding-3-small accepts up to 8,191 tokens per input string. This is the model's context window for embedding.

Key details:

  • Tokenizer: cl100k_base (same as GPT-4)
  • 1 token ≈ 4 characters in English, ≈ 0.75 words
  • 8,191 tokens ≈ 6,100 words ≈ 24,000 characters
  • Text exceeding 8,191 tokens is truncated (not rejected)

How to Count Tokens Before Sending#

Handling Long Documents#

For documents longer than 8,191 tokens, you have two options:

  1. Chunking: Split into overlapping segments and embed each one
  2. Summarize first: Use an LLM to summarize, then embed the summary

How to Use the Text-Embedding-3-Small API#

Python (OpenAI SDK)#

Batch Embedding (Multiple Texts)#

Send up to 2,048 texts in a single request for better throughput:

cURL#

Response Format#

Base64 Output Format#

For bandwidth-sensitive applications, request base64-encoded output:

Error Handling#

Text-Embedding-3-Small vs Text-Embedding-3-Large#

This is the most common comparison. Here's the full breakdown:

Featuretext-embedding-3-smalltext-embedding-3-large
Default dimensions15363072
Adjustable dimensions256 – 1536256 – 3072
Max tokens8,1918,191
MTEB score62.364.6
Price / 1M tokens$0.020$0.130
Price via Crazyrouter$0.016$0.100
Relative qualityBaseline+3.7% better
Relative costBaseline6.5x more expensive
Storage (float32)6 KB/vector12 KB/vector

When to Choose text-embedding-3-small#

  • Budget-conscious projects
  • High-volume embedding workloads (millions of documents)
  • RAG applications where "good enough" retrieval is fine
  • Prototyping and development
  • Applications where latency matters (smaller vectors = faster similarity search)

When to Choose text-embedding-3-large#

  • Search quality is the top priority and budget allows it
  • Legal, medical, or financial domains where precision matters
  • Small document collections where the 6.5x cost difference is negligible
  • You can use dimension reduction (e.g., 1024 dims) to get large-model quality at reduced storage

The Dimension Reduction Trick#

text-embedding-3-large at 1024 dimensions often outperforms text-embedding-3-small at 1536 dimensions — while using less storage:

Text-Embedding-3-Small vs Ada-002#

text-embedding-ada-002 was OpenAI's previous generation embedding model. Here's why you should migrate:

Featuretext-embedding-3-smalltext-embedding-ada-002
MTEB score62.361.0
Price / 1M tokens$0.020$0.100
Dimensions1536 (adjustable)1536 (fixed)
Dimension reduction✅ Yes❌ No
Multilingual (MIRACL)44.031.4

text-embedding-3-small is 5x cheaper, higher quality, and supports dimension reduction. There's no reason to stay on ada-002.

Migration note: Embeddings from different models are not compatible. Switching requires re-embedding all your documents.

Text-Embedding-3-Small Multilingual Support#

text-embedding-3-small supports multilingual text natively. On the MIRACL multilingual benchmark, it scores 44.0 — a massive improvement over ada-002's 31.4.

Supported languages include English, Chinese, Japanese, Korean, Spanish, French, German, Portuguese, Russian, Arabic, Hindi, and many more.

For applications that are primarily multilingual (100+ languages, cross-lingual retrieval at scale), also consider Cohere embed-v4 or the open-source BGE-M3.

Is Text-Embedding-3-Small Deprecated?#

No. As of 2026, text-embedding-3-small is fully active and supported by OpenAI. It is not deprecated and there is no announced deprecation date.

The model that is deprecated is the older text-embedding-ada-002. OpenAI recommends migrating from ada-002 to the text-embedding-3 series.

Timeline:

  • text-embedding-ada-002: Released December 2022, now legacy
  • text-embedding-3-small: Released January 2024, current recommended model
  • text-embedding-3-large: Released January 2024, current premium model

Building Semantic Search with Text-Embedding-3-Small#

Here's a complete working example:

Scaling with a Vector Database#

For production workloads with millions of documents, use a vector database instead of in-memory numpy:

Other popular vector databases that work well with text-embedding-3-small:

  • Pinecone: Fully managed, easiest to start
  • Weaviate: Open source, supports hybrid search
  • Qdrant: Open source, high performance
  • ChromaDB: Lightweight, great for prototyping
  • pgvector: PostgreSQL extension, no new infrastructure needed

Performance Benchmarks#

MTEB (Massive Text Embedding Benchmark)#

ModelAverageClassificationClusteringRetrieval
text-embedding-3-small62.367.141.251.7
text-embedding-3-large64.669.844.155.4
text-embedding-ada-00261.066.340.149.3
Cohere embed-v466.271.546.356.8
Voyage voyage-3-large67.172.047.258.1

Multilingual (MIRACL)#

ModelMIRACL Score
text-embedding-3-small44.0
text-embedding-3-large54.9
text-embedding-ada-00231.4

Latency#

text-embedding-3-small is one of the fastest commercial embedding models. Typical latency through Crazyrouter:

Batch SizeAvg Latency
1 text~50ms
10 texts~80ms
100 texts~200ms
1000 texts~800ms

Best Practices#

  1. Batch requests: Send multiple texts per API call (up to 2,048) to reduce overhead
  2. Cache embeddings: Never re-embed the same text — store results in a vector database or cache
  3. Normalize vectors: The API returns normalized vectors by default, but verify if using dimension reduction
  4. Choose dimensions wisely: Start with 1536, reduce only if storage or latency is a real constraint
  5. Use consistent models: Never mix embeddings from different models in the same index
  6. Chunk long documents: Split texts over 8,191 tokens with overlap for context continuity
  7. Monitor token usage: Track usage.total_tokens in responses to manage costs

Frequently Asked Questions#

What is text-embedding-3-small?#

text-embedding-3-small is OpenAI's cost-effective text embedding model. It converts text into 1536-dimensional numerical vectors that capture semantic meaning, enabling applications like semantic search, RAG, classification, and clustering.

How much does text-embedding-3-small cost?#

0.020per1milliontokensonOpenAIdirectly.Through[Crazyrouter](https://crazyrouter.com),itcosts0.020 per 1 million tokens on OpenAI directly. Through [Crazyrouter](https://crazyrouter.com), it costs 0.016 per 1M tokens — 20% cheaper. Embedding 10,000 pages of text costs roughly $0.15.

What is the token limit for text-embedding-3-small?#

8,191 tokens per input string. That's approximately 6,100 words or 24,000 characters in English. Text exceeding this limit is silently truncated.

How many dimensions does text-embedding-3-small output?#

1,536 dimensions by default. You can reduce this to any value between 256 and 1,536 using the dimensions parameter in the API request.

Is text-embedding-3-small deprecated?#

No. It is fully active and supported as of 2026. The deprecated model is the older text-embedding-ada-002.

What's the difference between text-embedding-3-small and text-embedding-3-large?#

text-embedding-3-large outputs 3,072 dimensions (vs 1,536), scores 3.7% higher on MTEB benchmarks, and costs 6.5x more (0.13vs0.13 vs 0.02 per 1M tokens). For most applications, the small model is sufficient.

Does text-embedding-3-small support multilingual text?#

Yes. It handles multiple languages natively and scores 44.0 on the MIRACL multilingual benchmark. No special configuration is needed — just pass text in any supported language.

Can I use text-embedding-3-small with LangChain?#

Yes. LangChain has built-in support:

Can I use text-embedding-3-small with LlamaIndex?#

Yes:

How does text-embedding-3-small compare to free/open-source models?#

Open-source models like BGE-M3 (MTEB 63.2) and E5-Mistral-7B (MTEB 66.6) can match or exceed its quality. The trade-off is you need GPU infrastructure to run them. text-embedding-3-small wins on convenience and total cost for small-to-medium workloads.

Summary#

text-embedding-3-small is the default choice for production embedding workloads in 2026. At 0.02/1Mtokens(or0.02/1M tokens (or 0.016 via Crazyrouter), it delivers strong quality across search, RAG, and classification — with the flexibility of dimension reduction and multilingual support.

Choose text-embedding-3-small when: you want the best price-to-performance ratio for embedding workloads.

Choose text-embedding-3-large when: you need maximum quality and the 6.5x cost increase is acceptable.

Access it through Crazyrouter for 20% lower pricing and a unified API that also gives you GPT-5, Claude, Gemini, and 300+ other models — all with one API key.