惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

B
Blog RSS Feed
T
The Blog of Author Tim Ferriss
P
Proofpoint News Feed
T
Threat Research - Cisco Blogs
T
The Exploit Database - CXSecurity.com
cs.AI updates on arXiv.org
cs.AI updates on arXiv.org
Application and Cybersecurity Blog
Application and Cybersecurity Blog
Threat Intelligence Blog | Flashpoint
Threat Intelligence Blog | Flashpoint
C
Cybersecurity and Infrastructure Security Agency CISA
D
Darknet – Hacking Tools, Hacker News & Cyber Security
Martin Fowler
Martin Fowler
GbyAI
GbyAI
P
Palo Alto Networks Blog
N
Netflix TechBlog - Medium
C
Cisco Blogs
Microsoft Security Blog
Microsoft Security Blog
G
Google Developers Blog
A
About on SuperTechFans
PCI Perspectives
PCI Perspectives
Scott Helme
Scott Helme
TaoSecurity Blog
TaoSecurity Blog
K
KPMG report finds enterprise disconnect between AI and its ROI | CIO
K
Kaspersky official blog
W
WeLiveSecurity
Y
Y Combinator Blog
Hacker News - Newest:
Hacker News - Newest: "LLM"
aimingoo的专栏
aimingoo的专栏
F
Fortinet All Blogs
有赞技术团队
有赞技术团队
人人都是产品经理
人人都是产品经理
月光博客
月光博客
N
News | PayPal Newsroom
Microsoft Azure Blog
Microsoft Azure Blog
G
GRAHAM CLULEY
爱范儿
爱范儿
The GitHub Blog
The GitHub Blog
MongoDB | Blog
MongoDB | Blog
V
V2EX
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻
Know Your Adversary
Know Your Adversary
博客园 - Franky
cs.CV updates on arXiv.org
cs.CV updates on arXiv.org
F
Full Disclosure
V
Vulnerabilities – Threatpost
V
Visual Studio Blog
Forbes - Security
Forbes - Security
Attack and Defense Labs
Attack and Defense Labs
MyScale Blog
MyScale Blog
Hacker News: Ask HN
Hacker News: Ask HN
T
Tor Project blog

DEV Community

Authentication Security Deep Dive: From Brute Force to Salted Hashing (With Java Examples) Why AI Systems Don’t Fail — They Drift Spilling beans for how i learn for exam😁"Reinforcement Learning Cheat Sheet" I Replaced Chrome with Safari for AI Browser Automation. Here's What Broke (and What Finally Worked) How Python Borrows Other People's Work The $40 Architecture: Processing 1 Billion API Requests with 99.99% Uptime Vibe Coding: A Workflow Guide (From Zero to SaaS) Most webhook security guides protect the wrong side. The scary part is delivery. Headless CMS for TanStack Start: Build a Blog with Cosmic EU Age Verification App "Hacked in 2 Minutes" — What Actually Happened Comfy Cloud’s delete function does not actually remove files Running AI Models on GPU Cloud Servers: A Beginner Guide Event-driven media intelligence with AWS Step Functions and Bedrock I scored 500 AI prompts across 8 quality dimensions — here's what broke How to Call Google Gemini API from Next.js (Free Tier, No Backend Needed) The Portal Protocol: Reclaiming Human Connection in the Age of AI How to Fix Your Team's Scattered Knowledge Problem With a Self-Hosted Forum Intro to tc Cloud Functors: A Graph-First Mental Model for the Modern Cloud Designing Multi-Tenant Backends With Both Ownership and Team Access I Built a Neumorphic CSS Library with 77+ Components — Here's What I Learned PostgreSQL Performance Optimization: Why Connection Pooling Is Critical at Scale Cómo construí un SaaS multi-rubro para gestionar expensas en Argentina con FastAPI + Vue 3 🚀 I Built an Ethical Hacking Scanner Tool – Open Source Project I Replaced /usage and /context in Claude Code With a Single Statusline A Pythonic Way to Handle Emails (IMAP/SMTP) with Auto-Discovery and AI-Ready Design I Collected 8.9 Million Polymarket Price Points — Here's What I Found About How Markets Really Move EcoTrack AI — Carbon Footprint Tracker & Dashboard Everyone's Using AI. No One Agrees How. 5 self-hosted ebook managers worth trying in 2026 Building Your First AI Agent with LangChain: From Chatbot to Autonomous Assistant Common SOC 2 Failures (Real World) Stop Vibe-Checking Your AI App: A Practical Guide to Evals How to Use SonarQube and SonarScanner Locally to Level Up Your Code Quality Your Next To-Do App Is Dead — I Replaced Mine with an OpenClaw AI Sign a Nostr event in 60 lines of Python using coincurve — no nostr-sdk, no nbxplorer, no rust toolchain ITGC Audit Explained Like You’re in Big 4 Patch Tuesday abril 2026: Microsoft parcha 163 vulnerabilidades y un zero-day en SharePoint Stop scraping everything: a better way to track competitor price changes Listing on MCPize + the Official MCP Registry while routing payments OUTSIDE the marketplace — how I kept 100% of my x402 revenue Building an AI-Powered Risk Intelligence System Using Serverless Architecture Why We Ripped Function Overloading Out of Our AI Toolchain Testing AI-Generated Code: How to Actually Know If It Works SaaS Churn Is Killing Your Business. Here Is What to Do About It (Without a Support Team) The Speed of AI Is No Longer Linear - And Self-Improving Models Are Why How to Implement RBAC for MCP Tools: A Practical Guide for Engineering Teams From Standard Quote to Persuasive Proposal: AI Automation for Arborists I built a CLI that scaffolds complete multi-tenant SaaS apps Axios CVE-2025–62718: The Silent SSRF Bug That Could Be Hiding in Your Node.js App Right Now The dashboard that ended our friendship Data Pipelines Explained Simply (and How to Build Them with Python) The Hidden Cost of AI Systems Nobody Talks About. undefined vs undeclared, and how typeof behaves Switching from file-based jobs to NATS/Kafka in Rust without changing code io_uring Adventures: Rust Servers That Love Syscalls Why Agentic AI is Killing the Traditional Database The POUR principles of web accessibility for developers and designers Quantum Neural Network 3D — A Deep Dive into Interactive WebGL Visualization How To Install Caveman In Codex On macOS And Windows Automation Pipeline Reliability: Why Your Workflow Breaks When Nobody Is Watching I Built an 'Open World' AI Coding Agent — It Works From ANY Folder From Freelancing to Product: A Tech Service Company's SaaS Transformation China's AI Giants: Adding Tencent Hunyuan & ByteDance Doubao to AI University (74 Providers) On the Vibe Coders and Their Lies clerk: Auto-Summarize Your Claude Code Sessions AI Weekly — 2026/04/10–04/17 | The Model Lockdown Is Here, but the Toolchain Is the Real Battleground AI 週報 — 2026/04/10–2026/04/17 模型封鎖潮來了,但工具鏈才是真戰場 Maybe this is how Open-Source apps are born... 🚀 Fine-Tune LLMs with LoRA and QLoRA: 2026 Guide tRPC v11 + Next.js App Router: End-to-End Type Safety Without the Boilerplate ShadCN UI in 2026: Why I Stopped Installing Component Libraries and Started Owning My Components SaaS Billing in React Server Components: Stripe + Supabase Without a Single `useEffect` Join our DEV Weekend Challenge — $1,000 in Prizes Across TEN winners! Submissions Due April 20 at 6:59 AM UTC. Implementing FSRS Spaced Repetition in Flutter + Supabase — Adding Memory Science to an AI Learning App "I Texted My Localhost From the Train — Claude Code Fixed the Bug Before I Got Home" I Built a Sales Prep AI and It Went Deeper Than Expected Design to Code #2: One JSON, Eleven Outputs Solving the 100M-Row Problem: A Summary Table Pattern for High-Volume Push Notification Logs Flutter Web With Wasm: What Actually Changes For Developers I Built 50 Royalty-Free Soundtracks for My Side Project in a Weekend Using AI Music Generation The Vibe Coding Security Checklist: 7 Things to Check Before You Ship Stop Letting Googlebot Guess Fix Your React App's SEO Right Desconstruindo o Streaming do LinkedIn: Como Criar um Engine de Extração de Vídeo de Alta Performance com HLS e FFmpeg (EDA Part-1) EDA (Exploratory Data Analysis) Explained With Real Life — Why Looking at Your Data Is the Most Important Step in Machine Learning Brand Relationship Management at Scale: Our 4-Touch Outreach System for 200+ Brands Why String.fromEnvironment() Might Return an Empty String in Dart JGuardrails 1.0.0 — Hardening Java LLM Apps Against Jailbreaks, Toxicity, and Prompt Injection Plan and Schedule a Full Week of Threads Content From One Claude Conversation Coding Cat Oran Ep3, Five Tables Changed Everything Updated: BFF Pattern I'm done watching freelancers get buried by 200 proposals. So I'm building the alternative. This is my first post BFS Algorithm in Java Step by Step Tutorial with Examples Tracking LLM Pricing Monthly: An Open Dataset for 22 AI Models How We Measure Content ROI on a Comparison Site: Revenue Attribution Without Perfect Data Introducing Nova AI Ops: The AI-Native Operating System for SRE Teams I built a free desktop video downloader for Windows — Grabbit How Talkie OCR Helps Vision-Impaired & Dyslexic Users Read the World Around Them VRCFaceTracking安装和iPhone面捕配置教程,有bug Even CrowdStrike Can't See Your Agents The Automation Gold Rush: What n8n Workflows and Claude Are Opening Up for Developers Right Now
Tokenization under the hood: BPE, WordPiece, SentencePiece, and Unigram compared
Tech_Nuggets · 2026-06-17 · via DEV Community

Tokenization under the hood: BPE, WordPiece, SentencePiece, and Unigram compared

You deploy a chatbot. English queries average 42 tokens each. Then a Spanish-speaking user sends "¿Cómo puedo restablecer mi contraseña?" and it eats 103 tokens. Two weeks later, the same model starts outputting "Ġcon" at the edges of its generations and you cannot tell if it is a bug or a feature. The finance team flags a 40% month-over-month cost increase that no one can explain.

This is what happens when tokenization is treated as invisible plumbing. Every major LLM pipeline uses one of four subword tokenization algorithms, and the choice determines vocabulary size, handling of rare words, cross-language efficiency, and inference cost. Understanding which one your model uses -- and why -- is the difference between shipping a cost-efficient product and discovering mid-quarter that your token-per-query ratio quietly doubled.

Why this matters

Tokenization directly controls three things that hit your bottom line:

Inference cost. LLM APIs charge by token. A model using a 32K-vocab BPE tokenizer may break "restablecer" into 8 tokens, while a 100K-vocab Unigram tokenizer handles it in 3. Over a million queries, that difference adds up to real money.

Vocabulary coverage. Rare words, code syntax, and multilingual text stress the tokenizer. A poorly fitting vocabulary means longer sequences, which means slower generation and higher cost.

Model behavior. The tokenizer is the model's entire view of language. If your tokenizer encodes "cowboy" as ["cow", "boy"], the model learns something different than if it encodes it as ["c", "owb", "oy"]. This affects everything from spelling ability to cross-lingual transfer.

The four tokenization algorithms

Every modern tokenizer takes raw text, optionally pre-tokenizes it into words (splitting on whitespace and punctuation), then breaks words into subword units from a fixed-size vocabulary. The difference is in how that vocabulary is built and how segmentation decisions are made.

1. BPE (Byte-Pair Encoding)

BPE was introduced in 1994 for data compression and adapted for neural machine translation by Sennrich et al. in 2016. OpenAI adopted it for GPT-2 and it remains the core of GPT-4o, Llama 3, and most modern LLMs.

How it works: Start with every individual character as a token. Count all adjacent token pairs, merge the most frequent pair into a new token, add it to the vocabulary, and repeat until you hit the target vocabulary size.

Vocabulary size goal: 16
Initial vocabulary: [a, b, c, d, e, f, g, h, i, j, k, l, m, n, o, p, q, r, s, t, u, v, w, x, y, z,  , ., ,]
Training corpus: "low low low low low low low low lower lowest lowest lowest lowest lowest lowest lowest"

Step 1: Count pairs -> ("l", "o") appears 30 times, merge -> "lo"
Step 2: Count pairs -> ("lo", "w") appears 20 times, merge -> "low"
Step 3: Count pairs -> ("low", "e") appears 10 times, merge -> "lowe"
Step 4: Count pairs -> ("lowe", "r") appears 4 times, merge -> "lower"
Step 5: Count pairs -> ("low", "e") appears 6 times... wait, "low"+"e" appears
         in "lowest" fragments, merge -> "lowe" already exists, so merge "lowe"+"st"
...

BPE is greedy and deterministic: for any input, the segmentation is the same every time. The algorithm applies the learned merge rules in order. OpenAI's GPT-4o uses o200k_base (200,096 tokens), GPT-4 used cl100k_base (100,256 tokens), and GPT-2 used a 50,257-token vocabulary.

Who uses it: GPT-4o, GPT-4, GPT-3.5, Llama 2, Llama 3 (via SentencePiece), DeepSeek, Mistral.

2. WordPiece

Google introduced WordPiece for Japanese/Korean voice search in 2012, and it powered BERT in 2018. It is often described as "BPE but with likelihood instead of frequency."

How it works: The algorithm starts the same way as BPE -- character-level initial tokens -- but instead of counting raw frequencies, it merges the pair that maximizes the likelihood of the training data under the current vocabulary. In practice this means it picks the pair whose merge increases the corpus-likelihood the most.

Compare merge candidates:
  Merge ("a", "b") -> new token likelihood gain: 0.0032
  Merge ("th", "e") -> new token likelihood gain: 0.0417
  Merge ("ing", " ") -> new token likelihood gain: 0.0281

WordPiece picks ("th", "e") because the probability lift is largest.

The result is that WordPiece tends to create tokens that are more linguistically meaningful -- common prefixes, suffixes, and root words -- compared to BPE's purely frequency-driven merges.

Who uses it: BERT, DistilBERT, ELECTRA, and most encoder-only models from Google.

3. SentencePiece

SentencePiece is a framework by Google (Kudo and Richardson, 2018) that wraps both BPE and Unigram tokenization. Its defining innovation: it operates directly on raw text without requiring a pre-tokenization step. Most tokenizers need whitespace/punctuation splitting before training, which ties them to a language-specific concept of "word." SentencePiece treats the input as a raw Unicode byte sequence, making it truly language-agnostic.

Raw text: "Hello世界"
With pre-tokenization: ["Hello", "世界"]  <- language-dependent
SentencePiece raw: "H", "e", "l", "l", "o", "世", "界"  <- no pre-tokenization needed

Who uses it: Llama 2, Llama 3, Gemma, T5, XLNet (in Unigram mode).

4. Unigram Language Model

Unigram (Kudo, 2018) flips the problem around. Instead of greedily building up a vocabulary from characters, it starts with a large vocabulary of candidate tokens and prunes it down using a probabilistic model.

How it works: Unigram models each token as an independent event and learns a probability distribution over the vocabulary. The segmentation of a word is the sequence of tokens whose probabilities multiply to the highest score.

Vocabulary: {"UN": 0.02, "UNIC": 0.005, "NI": 0.01, "UNI": 0.015, ...}

Input: "UNICORN"
Candidate segmentations and their scores:
  UN + I + C + O + R + N  -> 0.02 * 0.03 * 0.04 * 0.02 * 0.01 * 0.02 = 1.92e-12
  UNI + C + O + R + N     -> 0.015 * 0.04 * 0.02 * 0.01 * 0.02 = 2.4e-10
  UNIC + O + R + N        -> 0.005 * 0.02 * 0.01 * 0.02 = 2.0e-9  <-- best

Unigram picks the highest-probability segmentation: UNIC + O + R + N

Because Unigram evaluates multiple candidate segmentations and chooses the best one probabilistically, it is slower to tokenize than BPE but produces more consistent token-to-meaning mappings. The probabilistic nature also enables subword regularization -- randomly sampling alternative segmentations during training to improve robustness.

Who uses it: T5, XLNet, ALBERT, and SentencePiece in Unigram mode.

Algorithm comparison

Property BPE WordPiece SentencePiece (BPE) Unigram LM
Vocabulary building Greedy merge by frequency Greedy merge by likelihood Greedy merge by frequency (same as BPE) Start big, prune by likelihood
Pre-tokenization required Yes (whitespace/punctuation) Yes No (raw bytes) No (raw bytes)
Deterministic segmentation Yes Yes Yes No (sampling possible)
Typical vocab size 32K-200K 30K 32K-128K 32K-256K
Speed Fast Fast Fast Medium (Viterbi decoding)
Multilingual handling Weak (needs large vocab) Moderate Best (byte-level) Best (byte-level + sampling)
Rare word handling Decomposes to chars Decomposes to chars Decomposes to bytes Decomposes to subwords
Primary users OpenAI, Meta, Mistral Google (BERT) Meta (Llama), Google (Gemma) Google (T5, XLNet)

What this looks like in practice

Here is a Python snippet using tiktoken (OpenAI's BPE tokenizer library) to see how different inputs break apart:

import tiktoken

# GPT-4o uses o200k_base encoding
enc = tiktoken.get_encoding("o200k_base")

test_strings = [
    "Hello, world!",
    "restablecer",          # Spanish
    "Das ist fantastisch",  # German
    "こんにちは",            # Japanese
    "def fibonacci(n): return n if n <= 1 else fibonacci(n-1) + fibonacci(n-2)",
]

for s in test_strings:
    tokens = enc.encode(s)
    token_strs = [enc.decode([t]) for t in tokens]
    print(f"{s!r:45s} -> {len(tokens):3d} tokens: {token_strs[:6]}...")

Output (approximate for o200k_base):

Hello, world!                                    -> 3 tokens: ['Hello', ',', ' world']
restablecer                                      -> 8 tokens: ['rest', 'able', 'cer', ...]
Das ist fantastisch                              -> 6 tokens: ['Das', ' ist', ' fant', 'ast', 'isch', ...]
こんにちは                                         -> 5 tokens: ['こ', 'ん', 'に', 'ち', 'は']
def fibonacci(n): return n if n <= 1 else ...   -> 22 tokens: ['def', ' fib', 'onacci', ...]

Notice how the Spanish word takes 8 tokens while an analogous English word of similar length might take 3-4. This is the cost asymmetry that shows up on your monthly bill.

Here is a diagram showing how a single word passes through each tokenizer type:

flowchart TD
    A["Input: 'unbelievable'"] --> B["Pre-tokenization<br/>(split on space/punct)"]
    B --> C{"Tokenizer type?"}

    C -->|BPE| D["Lookup in vocab: 'un' + 'believable'<br/>If 'believable' not found:<br/>'b' + 'el' + 'ievable' ...<br/>Greedy character-level fallback"]
    C -->|WordPiece| E["Lookup longest prefix: 'un'<br/>Try '##believable'<br/>If not found: '##b' + '##el' + ...<br/>Likelihood-based merging"]
    C -->|SentencePiece| F["Byte-level segmentation<br/>No pre-tokenization<br/>BPE merge rules on raw bytes<br/>'un' + 'bel' + 'ievable'"]
    C -->|Unigram| G["Score all candidate segmentations<br/>Pick highest-probability path<br/>'un' + 'believ' + 'able'<br/>Probabilistic, may vary"]

    D --> H["Output tokens"]
    E --> H
    F --> H
    G --> H

Common pitfalls

Assuming all tokenizers handle multilingual text equally. BPE-based tokenizers that rely on space-prefix pre-tokenization (like cl100k_base) degrade significantly on CJK and Indic scripts where whitespace does not separate words. SentencePiece models handle these better because they operate at the byte level. If your user base spans non-Latin scripts, check your tokenizer's cross-language efficiency before picking a model.

Tying your prompt design to the wrong encoding. An instruction like "Output the result as JSON" costs 5 tokens with cl100k_base but 7 tokens with o200k_base. Developers who craft prompts for GPT-4 and then migrate to a model with a different tokenizer silently change the prompt's token boundary handoff, which can shift output quality.

Ignoring the tokenizer's role in fine-tuning. When you fine-tune a model, you can extend the vocabulary -- but doing so requires initializing new embedding vectors, and the model will behave unpredictably with the new tokens for the first few thousand steps. Most practitioners are better off using the existing vocabulary and handling out-of-vocabulary tokens via character-level fallback.

The "split on prefix space" trap. Most BPE tokenizers add a space before each word during pre-tokenization (byte-pair encoding operates on the string " Hello" not "Hello"). This means "Hello" (capitalized, start of sentence) and "hello" (lowercase, mid-sentence) share the same token " Hello" if the space prefix is consistent. But if your text formatting changes -- removing trailing spaces, using non-standard punctuation -- you can tokenize the same semantic content into dramatically more tokens.

Forgetting that tokenizer version matters. p50k_base and cl100k_base and o200k_base all use BPE with different pre-tokenization rules and vocab sizes. A comparison of two models' outputs is meaningless if you used different tokenizers to count their tokens. Pin your tiktoken version (tiktoken==0.13.0 as of June 2026) and your encoding name in every evaluation script.

When NOT to use it

When you need exact character-level control. Tokenization destroys alignment between text characters and model internals. If you are building a spelling corrector, a character-level model (like ByT5 or CANINE) produces better results than any subword tokenizer.

When latency is the absolute priority. SentencePiece Unigram and WordPiece both require running a language model or Viterbi decoder to segment text. BPE is simpler and faster. If you are measuring single-digit millisecond TTFT budgets, use a pure BPE tokenizer and keep the vocabulary under 50K.

When you are building a single-language, domain-specific model. If your entire task is English medical text classification, you can build a custom BPE vocabulary (15K-20K tokens) that outperforms the general-purpose 100K vocabulary in both speed and perplexity. The general vocabularies are optimized for web-scale diversity, not domain density.

When you need reversible tokenization. Subword tokenization is lossy. You cannot reconstruct the original string perfectly from the token IDs if the tokenizer applied normalization (lowercasing, NFKC Unicode normalization, etc.). If you need byte-level round-trips, use a byte-level tokenizer (like the one in ByT5 or CANINE).

When you are benchmarking across model families. Comparing GPT-4o (200K vocab, BPE) against Llama 3 (32K vocab, SentencePiece BPE) by token count is comparing apples to oranges. Always benchmark on character or byte cost, not token cost, when models use different tokenizers.

TL;DR

  • BPE (GPT-4o, Llama 3, Mistral) builds vocabulary by merging the most frequent character pairs greedily. Deterministic, fast, but weak on multilingual text.
  • WordPiece (BERT, ELECTRA) merges by likelihood gain rather than frequency. Produces more linguistically meaningful tokens but requires pre-tokenization.
  • SentencePiece (Llama 3, Gemma, T5) wraps BPE and Unigram, operating on raw bytes without pre-tokenization. Best multilingual handling.
  • Unigram (T5, XLNet) starts with a large vocabulary and prunes it by likelihood. Supports subword regularization and produces more consistent token-to-meaning alignments at the cost of slower segmentation.
  • Tokenizer choice directly impacts inference cost: a 32K vocab English-optimized tokenizer and a 200K vocab general tokenizer will produce very different token counts for the same multilingual input.
  • Pin your tokenizer version and encoding name when reporting any token-count metric. Differences between cl100k_base and o200k_base can shift token counts by 15-30% on the same text.

Next post

When you know which tokenizer your model uses, the next question is how to prepare your data so that tokenizer wastes as few tokens as possible. That means strategic prompt design, choosing the right model for your language mix, and building evaluation pipelines that measure token efficiency alongside accuracy. We will cover token-efficient prompt engineering in the next post -- including a concrete method for estimating your per-user token consumption before you deploy.