惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

博客园 - 叶小钗
Microsoft Azure Blog
Microsoft Azure Blog
Stack Overflow Blog
Stack Overflow Blog
Jina AI
Jina AI
Vercel News
Vercel News
H
Help Net Security
Martin Fowler
Martin Fowler
美团技术团队
云风的 BLOG
云风的 BLOG
Y
Y Combinator Blog
阮一峰的网络日志
阮一峰的网络日志
MyScale Blog
MyScale Blog
钛媒体:引领未来商业与生活新知
钛媒体:引领未来商业与生活新知
博客园 - 三生石上(FineUI控件)
博客园 - 司徒正美
人人都是产品经理
人人都是产品经理
Engineering at Meta
Engineering at Meta
G
Google Developers Blog
Blog — PlanetScale
Blog — PlanetScale
MongoDB | Blog
MongoDB | Blog
宝玉的分享
宝玉的分享
小众软件
小众软件
T
Tailwind CSS Blog
WordPress大学
WordPress大学

DEV Community

Authentication Security Deep Dive: From Brute Force to Salted Hashing (With Java Examples) Why AI Systems Don’t Fail — They Drift Spilling beans for how i learn for exam😁"Reinforcement Learning Cheat Sheet" I Replaced Chrome with Safari for AI Browser Automation. Here's What Broke (and What Finally Worked) How Python Borrows Other People's Work The $40 Architecture: Processing 1 Billion API Requests with 99.99% Uptime Vibe Coding: A Workflow Guide (From Zero to SaaS) Most webhook security guides protect the wrong side. The scary part is delivery. Headless CMS for TanStack Start: Build a Blog with Cosmic EU Age Verification App "Hacked in 2 Minutes" — What Actually Happened Comfy Cloud’s delete function does not actually remove files Running AI Models on GPU Cloud Servers: A Beginner Guide Event-driven media intelligence with AWS Step Functions and Bedrock I scored 500 AI prompts across 8 quality dimensions — here's what broke How to Call Google Gemini API from Next.js (Free Tier, No Backend Needed) The Portal Protocol: Reclaiming Human Connection in the Age of AI How to Fix Your Team's Scattered Knowledge Problem With a Self-Hosted Forum Intro to tc Cloud Functors: A Graph-First Mental Model for the Modern Cloud Designing Multi-Tenant Backends With Both Ownership and Team Access I Built a Neumorphic CSS Library with 77+ Components — Here's What I Learned PostgreSQL Performance Optimization: Why Connection Pooling Is Critical at Scale Cómo construí un SaaS multi-rubro para gestionar expensas en Argentina con FastAPI + Vue 3 🚀 I Built an Ethical Hacking Scanner Tool – Open Source Project I Replaced /usage and /context in Claude Code With a Single Statusline A Pythonic Way to Handle Emails (IMAP/SMTP) with Auto-Discovery and AI-Ready Design I Collected 8.9 Million Polymarket Price Points — Here's What I Found About How Markets Really Move EcoTrack AI — Carbon Footprint Tracker & Dashboard Everyone's Using AI. No One Agrees How. 5 self-hosted ebook managers worth trying in 2026 Building Your First AI Agent with LangChain: From Chatbot to Autonomous Assistant
RAG- Understanding of Embedding
Ramya Peruma · 2026-05-18 · via DEV Community

Ramya Perumal

What is Embedding?

After text is split into chunks, the next process is called embedding. In this step, each chunk is converted into vectors (points in vector space). In vector-based RAG systems, chunks are converted into vectors so that semantic search can be performed efficiently.

Why Do We Need to Convert Chunks into Vectors?

The main goal of a RAG application is to achieve semantic search.

Semantic
Example
The word feline is related to the cat family, even though the words are different. Understanding that “feline” and “cat” are related is called semantic understanding.

Similarity
When a user asks a query, semantically related chunks are returned even though the exact words in the chunks may be different.

Semantic Similarity
Semantic similarity combines:

  • Intent
  • Context
  • Meaning

The purpose is to establish relationships between the user query and the documents stored in the RAG system. This allows the system to retrieve relevant information from the database and provide it to the LLM for further processing.

Words that are semantically related are usually stored closer together in multi-dimensional vector space.

Cosine Similarity
To determine how close vectors are to each other, cosine similarity is commonly used.

When a user query arrives:

  1. The query is converted into a vector
  2. Cosine similarity is calculated between the query vector and stored vectors
  3. The closest vectors are retrieved

Retrieval Methodologies

Two major retrieval methodologies are used:

1. KNN (K-Nearest Neighbors)
KNN compares the query vector with all stored vectors one by one to find the nearest neighbors.

Advantage
More accurate retrieval

Disadvantage
Slow for very large datasets

2. ANN (Approximate Nearest Neighbors)
ANN approximately finds the nearest vectors instead of comparing every single point.

This method is mainly used when:

  • The document volume is huge
  • Faster retrieval is required
  • Time constraints exist

ANN improves retrieval speed while sacrificing a small amount of accuracy.

Why Cosine Similarity Instead of Sine or Tangent?

Cosine similarity works effectively because:
If two vectors are very close and highly related, the cosine similarity value approaches 1. If the angle between vectors increases, the cosine similarity value decreases, meaning the vectors are less related

Why Not Sine or Tangent?
For small angles:

  • Sine values remain close to 0
  • Tangent values can fluctuate significantly

These measurements are not stable for semantic comparison. Cosine similarity provides a more reliable way to measure semantic closeness between vectors.

Embedding Dimensions
Embedding models can generate vectors with dimensions ranging from 256 to 3000 or more.

The dimension size depends on the embedding model and the amount of contextual information it captures.

Generally:

  • Higher dimensions capture richer semantic information
  • Lower dimensions are faster and cheaper but may lose context

Types of Embedding Models
Choosing an embedding model completely depends on the application scenario.

1. Based on Query Type

Symmetric Models
Symmetric embedding models are used when the query and the documents are similar in structure and length.

Examples
nomic-embed-text
Qwen embeddings

These are commonly used in semantic search systems.

Asymmetric Models
Asymmetric embedding models are used when:

  • Queries are short
  • Documents are long

Example
Google Gemini embedding models

These models are optimized for retrieving long documents from short queries.

2. Based on Retrieval Type

Dense Embeddings
Dense embeddings mainly focus on semantic meaning.

These embeddings generate dense vectors where most values contain meaningful information.

Examples
Cohere embedding models
ChatGPT OSS 120B embeddings

Advantage
Better semantic understanding

Sparse Embeddings
Sparse embeddings mainly focus on exact keyword matching.

They commonly use the BM25 (Best Match 25) algorithm, which is based on:

  • TF (Term Frequency)
  • IDF (Inverse Document Frequency)

TF-IDF Concepts

TF (Term Frequency)
Measures how many times a word appears in a document.

IDF (Inverse Document Frequency)
Measures how important a word is across the entire document collection.Words that appear too frequently across all documents are considered less important.

Transformer Architecture

The transformer architecture was a major breakthrough for LLMs.
Transformers mainly contain:

  • Encoder
  • Decoder

Encoder
The encoder converts text into embeddings (vectors).

Decoder
The decoder converts embeddings back into human-readable text after processing.

This architecture enables modern LLMs to understand and generate natural language effectively.

Choosing a Vector Database

Chroma
Open source
Easy to set up
Suitable for basic and small-scale applications

FAISS
Better for large document collections
Optimized for high-performance semantic search
Commonly used in production-scale retrieval systems