惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

Hacker News: Ask HN
Hacker News: Ask HN
Recent Commits to openclaw:main
Recent Commits to openclaw:main
cs.CV updates on arXiv.org
cs.CV updates on arXiv.org
C
Check Point Blog
S
Security Affairs
Hacker News - Newest:
Hacker News - Newest: "LLM"
S
Secure Thoughts
Recorded Future
Recorded Future
cs.AI updates on arXiv.org
cs.AI updates on arXiv.org
T
The Blog of Author Tim Ferriss
B
Blog
C
Cybersecurity and Infrastructure Security Agency CISA
Google DeepMind News
Google DeepMind News
CTFtime.org: upcoming CTF events
CTFtime.org: upcoming CTF events
A
Arctic Wolf
T
The Exploit Database - CXSecurity.com
Stack Overflow Blog
Stack Overflow Blog
T
Threat Research - Cisco Blogs
GbyAI
GbyAI
AWS News Blog
AWS News Blog
MongoDB | Blog
MongoDB | Blog
Y
Y Combinator Blog
Google Online Security Blog
Google Online Security Blog
T
Troy Hunt's Blog
I
InfoQ
L
LINUX DO - 热门话题
WordPress大学
WordPress大学
C
Cisco Blogs
G
GRAHAM CLULEY
The Register - Security
The Register - Security
A
About on SuperTechFans
让小产品的独立变现更简单 - ezindie.com
让小产品的独立变现更简单 - ezindie.com
Schneier on Security
Schneier on Security
Project Zero
Project Zero
H
Hackread – Cybersecurity News, Data Breaches, AI and More
P
Privacy & Cybersecurity Law Blog
Cloudbric
Cloudbric
H
Hacker News: Front Page
小众软件
小众软件
雷峰网
雷峰网
The Hacker News
The Hacker News
www.infosecurity-magazine.com
www.infosecurity-magazine.com
T
Tor Project blog
博客园 - 聂微东
N
Netflix TechBlog - Medium
V
Vulnerabilities – Threatpost
The GitHub Blog
The GitHub Blog
腾讯CDC
P
Palo Alto Networks Blog
Scott Helme
Scott Helme

Crazyrouter Blog

Gemini CLI Complete Guide 2026: Repo Automation, CI Agents, and Multi-Model Routing Ideogram AI Guide 2026: Brand Design Automation, API Workflows, and Alternatives GLM 4.6 API Guide 2026: Agents, RAG, Tool Calling, and Bilingual Apps WAN 2.2 Animate Tutorial 2026: Character Consistency, Shot Control, and API Workflows Google Veo3 API Guide 2026: Production Video Pipelines, Prompts, Pricing, and Fallbacks AI API Pricing Comparison 2026: Text, Image, Video, Caching, and Router Costs Codex CLI Installation Guide 2026: Windows, macOS, Linux, Proxies, and CI Setup How to Get a Claude API Key in 2026: Secure Setup for Teams, CI, and Alternatives Gemini Advanced Review 2026: Is It Worth It for Coding, Research, and API Teams? Claude Code Pricing Guide 2026: Team Agent Budgets, API Fallbacks, and Cost Control Seedance 2.0 Pricing: Convert 46 CNY per Million Tokens to Cost per Second Seedance 2.0 计费详解:46元/百万Token换算成每秒多少钱 Seedance 2.0料金解説:100万Tokenあたり46元を1秒あたりコストに換算 Gemini CLI 使用教程 2026:安装、代码示例、代理环境与 API 接入 Gemini 是什么?2026 完整介绍、API 使用教程与价格对比 Qwen2.5-Omni Guide 2026: Real-Time Voice, Vision, and Multimodal Agents Kimi K2 Thinking Guide 2026: Reasoning Workflows, Evals, and Cost Control Google Veo3 API Guide 2026: Batch Video Pipelines, Pricing, and Fallbacks Codex CLI Installation Guide 2026: macOS, Linux, WSL, Proxies, and Dev Containers How to Get a Claude API Key in 2026: Safe Production Setup and Alternatives AI API Pricing Comparison 2026: GPT, Claude, Gemini, Video, and Agent Workloads Gemini Advanced Review 2026: Is It Worth It for Developer Teams? Claude Code Pricing Guide 2026: API Fallbacks, Team Seats, and Budget Control Seedream 4.0 API Tutorial 2026: Batch Image Generation, Product Creative, and Pricing Qwen2.5-Omni Guide 2026: Real-Time Voice, Vision, Text Agents, and API Integration Kimi K2 Thinking Guide 2026: Reasoning Agents, Evaluation Workflows, and API Cost Control WAN 2.2 Animate Tutorial 2026: Character Motion, Shot Control, API Pipelines, and Pricing Google Veo3 API Guide 2026: Production Video Workflows, Prompts, Pricing, and Fallbacks AI API Pricing Comparison 2026: OpenAI, Claude, Gemini, DeepSeek, and Router Costs How to Get a Claude API Key in 2026: Setup, Security, Rotation, and Alternatives Codex CLI Installation Guide 2026: macOS, Linux, WSL, Proxies, and Devcontainers Gemini Advanced Review 2026: Is It Worth It for Developers and API Builders? Claude Code Pricing Guide 2026: CI Agents, Team Seats, and API Budget Planning 一個 API Key 呼叫 GPT、Claude、Gemini:5 分鐘設定教學 AI API Gateway for Singapore and Malaysia Developers: One Endpoint for GPT, Claude and Gemini AI API Gateway for Thai Developers: Use GPT, Claude and Gemini with One Key Cómo usar GPT, Claude y Gemini con una sola API key One API Key for GPT, Claude and Gemini: A Practical Setup for Central Asia Developers Gemini 3.5 Flash vs Claude レスポンスティアモデル:開発者はどちらを選ぶべきか Gemini 3.5 Flash vs Claude Response-Tier Models: Какую модель выбрать разработчику? Gemini 3.5 Flash vs Claude Response-Tier Models: Which One Should Developers Use? Gemini 3.5 Flash vs Gemini 3 Flash vs Gemini 2.5 Flash:実運用APIベンチマーク Gemini 3.5 Flash vs Gemini 3 Flash vs Gemini 2.5 Flash: Real API Benchmark text-embedding-3-large 值不值得用?和 text-embedding-3-small 的成本、效果与选型对比 用 text-embedding-3-large 搭建 RAG 知识库:从切块、向量化到检索排序 text-embedding-3-large 是干什么的?Embedding 模型入门与 RAG 场景详解 AI 扩图 API 指南 2026:Uncrop、Outpaint、gpt-image-2 和 Nano Banana 路线怎么选 How to Test Multiple AI Image Models with One API Key "How to Test Multiple AI Image Models with One API Key" Codex CLI Installation Guide: Setup on macOS, Linux, Windows WSL and CI/CD Gemini CLI 使用教程:开发者终端 AI 助手完全指南 Grok 4 免费使用教程:合法体验路径、API 接入与替代方案 Seedream 4.0 API Tutorial: ByteDance Image Generation for Production Pipelines Kimi K2 Thinking Model: Complete Developer Guide for Reasoning Workflows Luma Ray 2 Review: AI Video Generation Quality, Speed, and API Guide Pika 2.2 New Features Review: Scene Director, Sound Design, and API Updates Google Veo 3 API Guide: Video Generation with Audio for Developers AI Lip Sync Tools Comparison 2026: Best APIs for Talking Avatars and Video Dubbing Gemini Advanced Review May 2026: Is It Worth $20/Month for AI Power Users? Claude Code Pricing in May 2026: Max Plan, Opus 4, and Real Cost Breakdown Hermes Agent + Crazyrouter: One-Click Setup for 627+ AI Models Cursor 配置 Crazyrouter 教程:国内用上 GPT-5.4 / Claude 写代码 2026 年国内如何调用 Claude API?Claude Opus / Sonnet 接入完全指南 2026 年国内如何调用 GPT-5.4 API?完整接入指南(含代码示例) AI API 常见报错排查大全:401、429、500、timeout 一篇搞定 2026 年 AI API 中转站哪家好?六大平台横向对比评测 2026 年 DeepSeek R1 API 接入指南:国内最强推理模型怎么调用 Trình Tạo Meme & Sách Tô Màu Bằng AI Với GPT-image-2 — Những Dự Án Vui Mà Vẫn Kiếm Ra Tiền Dự Đoán Em Bé Tương Lai Bằng AI Với GPT-image-2 — Xem Con Bạn Có Thể Trông Như Thế Nào Chuyển Đổi Ảnh Sang Phong Cách Ghibli Với GPT-image-2 — Biến Mọi Bức Ảnh Thành Tranh Anime Tạo Mô Hình Nhân Vật Hành Động Bằng AI Với GPT-image-2 — Biến Bất Kỳ Ai Thành Đồ Chơi Trong Hộp GPT-image-2: Nhận Diện Khuôn Mặt Và Phân Tích Màu Sắc Bằng AI Xem chỉ tay với GPT-image-2 — Tạo bản phân tích chỉ tay chuyên nghiệp chỉ từ một bức ảnh GPT-image-2로 AI 밈 생성기 & 컬러링북 만들기 — 재미있고 수익도 되는 프로젝트 GPT-image-2로 AI 미래 아기 예측 — 우리 아이는 어떤 모습일까? GPT-image-2로 지브리 스타일 변환 — 사진을 애니메이션 아트로 바꾸기 GPT-image-2로 AI 액션 피규어 생성하기 — 누구나 박스형 피규어로 바꾸는 법 GPT-image-2로 AI 관상 분석 & 퍼스널 컬러 진단 — 두 가지 바이럴 활용법 완벽 가이드 GPT-image-2 실전 가이드:AI 손금 분석 — 손바닥 사진 한 장으로 전문 손금 인포그래픽 생성하기 GPT-image-2 で AI ミーム生成 & ぬりえブック制作 — 楽しくて本当に稼げるプロジェクト GPT-image-2 で AI 未来の赤ちゃん予測 — 将来の子どもの顔を見てみよう GPT-image-2 でジブリ風写真変換 — どんな写真もアニメアートに GPT-image-2 で AI アクションフィギュア生成 — 誰でもボックス入りおもちゃに変身 GPT-image-2 で AI 顔相診断 & パーソナルカラー分析 — 2つのバズ活用法を1本で解説 GPT-image-2 で AI 手相占い — 1枚の写真からプロ仕様の手相分析を生成 GPT-image-2 на практике: AI-генератор мемов и раскрасок — весёлые проекты, которые приносят деньги GPT-image-2 на практике: AI-предсказание будущего ребёнка — как будет выглядеть ваш малыш GPT-image-2 на практике: стиль Гибли — превратите любое фото в аниме-арт GPT-image-2 на практике: AI-генератор фигурок — превратите себя в коллекционную игрушку GPT-image-2 на практике: AI-физиогномика и анализ цветотипа — два вирусных кейса в одном гайде GPT-image-2 на практике: AI-хиромантия — генерация профессионального анализа ладони по фото GPT-image-2 实战:AI Meme 生成器 & 涂色书制作 — 好玩还能赚钱的两个项目 GPT-image-2 实战:AI 预测未来宝宝 — 看看你们的孩子长什么样 GPT-image-2 实战:吉卜力风格转换 — 把任何照片变成宫崎骏动画 GPT-image-2 实战:AI 手办生成器 — 把任何人变成盒装公仔 GPT-image-2 实战:AI 面相分析 & 个人色彩诊断 — 两大爆款玩法一文搞定 GPT-image-2 实战:AI 看手相 — 一张手掌照片生成专业手相分析图 AI Meme Generator & Coloring Book Creator with GPT-image-2 — Fun Projects That Actually Make Money AI Future Baby Prediction with GPT-image-2 — See What Your Child Might Look Like Ghibli Style Photo Transformation with GPT-image-2 — Turn Any Photo Into Anime Art
Text-Embedding-3-Small: Complete Guide to OpenAI's Most Popular Embedding Model (2026)
Crazyrouter Team · 2026-05-03 · via Crazyrouter Blog

text-embedding-3-small is OpenAI's cost-effective embedding model, released in January 2024. It converts text into 1536-dimensional vectors that capture semantic meaning — the foundation for semantic search, RAG pipelines, recommendation systems, and classification tasks.

This guide covers everything: pricing, token limits, dimensions, API usage, dimension reduction, performance benchmarks, and how it compares to text-embedding-3-large.

Text-Embedding-3-Small Quick Reference#

SpecValue
Model nametext-embedding-3-small
ProviderOpenAI
Default dimensions1536
Adjustable dimensions256 – 1536
Max input tokens8,191
Max batch size2,048 inputs per request
Pricing (OpenAI direct)$0.020 per 1M tokens
Pricing (Crazyrouter)$0.016 per 1M tokens
MTEB benchmark score62.3
MultilingualYes
Output formatfloat or base64
Release dateJanuary 25, 2024
StatusActive (not deprecated)

Text-Embedding-3-Small Pricing#

text-embedding-3-small costs 0.020per1milliontokens∗∗onOpenAIdirectly.Through[Crazyrouter](https://crazyrouter.com),thepricedropsto∗∗0.020 per 1 million tokens** on OpenAI directly. Through [Crazyrouter](https://crazyrouter.com), the price drops to **0.016 per 1M tokens — a 20% discount.

To put that in perspective:

Document VolumeApprox. TokensCost (OpenAI)Cost (Crazyrouter)
100 pages of text~75,000$0.0015$0.0012
10,000 pages~7.5M$0.15$0.12
1 million pages~750M$15.00$12.00
Wikipedia (English, full)~4.4B$88.00$70.40

Cost Comparison with Other Embedding Models#

ModelPrice / 1M TokensPrice via CrazyrouterDimensions
text-embedding-3-small$0.020$0.0161536
text-embedding-3-large$0.130$0.1003072
text-embedding-ada-002$0.1001536
Google text-embedding-005$0.00625$0.005768
Cohere embed-v4$0.1001024
Voyage voyage-3-large$0.1802048

text-embedding-3-small is 6.5x cheaper than text-embedding-3-large and 5x cheaper than the older ada-002 — while outperforming ada-002 on benchmarks.

For a deeper comparison of all embedding models, see our AI Embeddings Comparison 2026 Guide.

Text-Embedding-3-Small Dimensions#

The default output is a 1536-dimensional vector. But text-embedding-3-small supports dimension reduction via the dimensions parameter — you can request any value from 256 to 1536.

This is done using Matryoshka Representation Learning (MRL). The model is trained so that the first N dimensions of the vector carry the most important information. Truncating to fewer dimensions loses some nuance but keeps most of the semantic signal.

Dimension vs. Quality Tradeoff#

DimensionsMTEB ScoreStorage per VectorRelative Quality
1536 (default)62.36,144 bytes100%
1024~61.54,096 bytes~98.7%
768~60.83,072 bytes~97.6%
512~59.72,048 bytes~95.8%
256~57.81,024 bytes~92.8%

When to Reduce Dimensions#

  • 256 dimensions: Prototyping, low-resource environments, or when storage is the bottleneck
  • 512 dimensions: Good balance for mobile apps or edge deployments
  • 768 dimensions: Matches Google's embedding size — useful for migration
  • 1536 dimensions: Production workloads where quality matters most

Text-Embedding-3-Small Token Limit and Context Length#

text-embedding-3-small accepts up to 8,191 tokens per input string. This is the model's context window for embedding.

Key details:

  • Tokenizer: cl100k_base (same as GPT-4)
  • 1 token ≈ 4 characters in English, ≈ 0.75 words
  • 8,191 tokens ≈ 6,100 words ≈ 24,000 characters
  • Text exceeding 8,191 tokens is truncated (not rejected)

How to Count Tokens Before Sending#

Handling Long Documents#

For documents longer than 8,191 tokens, you have two options:

  1. Chunking: Split into overlapping segments and embed each one
  2. Summarize first: Use an LLM to summarize, then embed the summary

How to Use the Text-Embedding-3-Small API#

Python (OpenAI SDK)#

Batch Embedding (Multiple Texts)#

Send up to 2,048 texts in a single request for better throughput:

cURL#

Response Format#

Base64 Output Format#

For bandwidth-sensitive applications, request base64-encoded output:

Error Handling#

Text-Embedding-3-Small vs Text-Embedding-3-Large#

This is the most common comparison. Here's the full breakdown:

Featuretext-embedding-3-smalltext-embedding-3-large
Default dimensions15363072
Adjustable dimensions256 – 1536256 – 3072
Max tokens8,1918,191
MTEB score62.364.6
Price / 1M tokens$0.020$0.130
Price via Crazyrouter$0.016$0.100
Relative qualityBaseline+3.7% better
Relative costBaseline6.5x more expensive
Storage (float32)6 KB/vector12 KB/vector

When to Choose text-embedding-3-small#

  • Budget-conscious projects
  • High-volume embedding workloads (millions of documents)
  • RAG applications where "good enough" retrieval is fine
  • Prototyping and development
  • Applications where latency matters (smaller vectors = faster similarity search)

When to Choose text-embedding-3-large#

  • Search quality is the top priority and budget allows it
  • Legal, medical, or financial domains where precision matters
  • Small document collections where the 6.5x cost difference is negligible
  • You can use dimension reduction (e.g., 1024 dims) to get large-model quality at reduced storage

The Dimension Reduction Trick#

text-embedding-3-large at 1024 dimensions often outperforms text-embedding-3-small at 1536 dimensions — while using less storage:

Text-Embedding-3-Small vs Ada-002#

text-embedding-ada-002 was OpenAI's previous generation embedding model. Here's why you should migrate:

Featuretext-embedding-3-smalltext-embedding-ada-002
MTEB score62.361.0
Price / 1M tokens$0.020$0.100
Dimensions1536 (adjustable)1536 (fixed)
Dimension reduction✅ Yes❌ No
Multilingual (MIRACL)44.031.4

text-embedding-3-small is 5x cheaper, higher quality, and supports dimension reduction. There's no reason to stay on ada-002.

Migration note: Embeddings from different models are not compatible. Switching requires re-embedding all your documents.

Text-Embedding-3-Small Multilingual Support#

text-embedding-3-small supports multilingual text natively. On the MIRACL multilingual benchmark, it scores 44.0 — a massive improvement over ada-002's 31.4.

Supported languages include English, Chinese, Japanese, Korean, Spanish, French, German, Portuguese, Russian, Arabic, Hindi, and many more.

For applications that are primarily multilingual (100+ languages, cross-lingual retrieval at scale), also consider Cohere embed-v4 or the open-source BGE-M3.

Is Text-Embedding-3-Small Deprecated?#

No. As of 2026, text-embedding-3-small is fully active and supported by OpenAI. It is not deprecated and there is no announced deprecation date.

The model that is deprecated is the older text-embedding-ada-002. OpenAI recommends migrating from ada-002 to the text-embedding-3 series.

Timeline:

  • text-embedding-ada-002: Released December 2022, now legacy
  • text-embedding-3-small: Released January 2024, current recommended model
  • text-embedding-3-large: Released January 2024, current premium model

Building Semantic Search with Text-Embedding-3-Small#

Here's a complete working example:

Scaling with a Vector Database#

For production workloads with millions of documents, use a vector database instead of in-memory numpy:

Other popular vector databases that work well with text-embedding-3-small:

  • Pinecone: Fully managed, easiest to start
  • Weaviate: Open source, supports hybrid search
  • Qdrant: Open source, high performance
  • ChromaDB: Lightweight, great for prototyping
  • pgvector: PostgreSQL extension, no new infrastructure needed

Performance Benchmarks#

MTEB (Massive Text Embedding Benchmark)#

ModelAverageClassificationClusteringRetrieval
text-embedding-3-small62.367.141.251.7
text-embedding-3-large64.669.844.155.4
text-embedding-ada-00261.066.340.149.3
Cohere embed-v466.271.546.356.8
Voyage voyage-3-large67.172.047.258.1

Multilingual (MIRACL)#

ModelMIRACL Score
text-embedding-3-small44.0
text-embedding-3-large54.9
text-embedding-ada-00231.4

Latency#

text-embedding-3-small is one of the fastest commercial embedding models. Typical latency through Crazyrouter:

Batch SizeAvg Latency
1 text~50ms
10 texts~80ms
100 texts~200ms
1000 texts~800ms

Best Practices#

  1. Batch requests: Send multiple texts per API call (up to 2,048) to reduce overhead
  2. Cache embeddings: Never re-embed the same text — store results in a vector database or cache
  3. Normalize vectors: The API returns normalized vectors by default, but verify if using dimension reduction
  4. Choose dimensions wisely: Start with 1536, reduce only if storage or latency is a real constraint
  5. Use consistent models: Never mix embeddings from different models in the same index
  6. Chunk long documents: Split texts over 8,191 tokens with overlap for context continuity
  7. Monitor token usage: Track usage.total_tokens in responses to manage costs

Frequently Asked Questions#

What is text-embedding-3-small?#

text-embedding-3-small is OpenAI's cost-effective text embedding model. It converts text into 1536-dimensional numerical vectors that capture semantic meaning, enabling applications like semantic search, RAG, classification, and clustering.

How much does text-embedding-3-small cost?#

0.020per1milliontokensonOpenAIdirectly.Through[Crazyrouter](https://crazyrouter.com),itcosts0.020 per 1 million tokens on OpenAI directly. Through [Crazyrouter](https://crazyrouter.com), it costs 0.016 per 1M tokens — 20% cheaper. Embedding 10,000 pages of text costs roughly $0.15.

What is the token limit for text-embedding-3-small?#

8,191 tokens per input string. That's approximately 6,100 words or 24,000 characters in English. Text exceeding this limit is silently truncated.

How many dimensions does text-embedding-3-small output?#

1,536 dimensions by default. You can reduce this to any value between 256 and 1,536 using the dimensions parameter in the API request.

Is text-embedding-3-small deprecated?#

No. It is fully active and supported as of 2026. The deprecated model is the older text-embedding-ada-002.

What's the difference between text-embedding-3-small and text-embedding-3-large?#

text-embedding-3-large outputs 3,072 dimensions (vs 1,536), scores 3.7% higher on MTEB benchmarks, and costs 6.5x more (0.13vs0.13 vs 0.02 per 1M tokens). For most applications, the small model is sufficient.

Does text-embedding-3-small support multilingual text?#

Yes. It handles multiple languages natively and scores 44.0 on the MIRACL multilingual benchmark. No special configuration is needed — just pass text in any supported language.

Can I use text-embedding-3-small with LangChain?#

Yes. LangChain has built-in support:

Can I use text-embedding-3-small with LlamaIndex?#

Yes:

How does text-embedding-3-small compare to free/open-source models?#

Open-source models like BGE-M3 (MTEB 63.2) and E5-Mistral-7B (MTEB 66.6) can match or exceed its quality. The trade-off is you need GPU infrastructure to run them. text-embedding-3-small wins on convenience and total cost for small-to-medium workloads.

Summary#

text-embedding-3-small is the default choice for production embedding workloads in 2026. At 0.02/1Mtokens(or0.02/1M tokens (or 0.016 via Crazyrouter), it delivers strong quality across search, RAG, and classification — with the flexibility of dimension reduction and multilingual support.

Choose text-embedding-3-small when: you want the best price-to-performance ratio for embedding workloads.

Choose text-embedding-3-large when: you need maximum quality and the 6.5x cost increase is acceptable.

Access it through Crazyrouter for 20% lower pricing and a unified API that also gives you GPT-5, Claude, Gemini, and 300+ other models — all with one API key.