惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

V
Visual Studio Blog
Y
Y Combinator Blog
奇客Solidot–传递最新科技情报
奇客Solidot–传递最新科技情报
Hugging Face - Blog
Hugging Face - Blog
钛媒体:引领未来商业与生活新知
钛媒体:引领未来商业与生活新知
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻
The Cloudflare Blog
L
LangChain Blog
美团技术团队
N
Netflix TechBlog - Medium
量子位
酷 壳 – CoolShell
酷 壳 – CoolShell
B
Blog
博客园 - 司徒正美
爱范儿
爱范儿
D
DataBreaches.Net
月光博客
月光博客
U
Unit 42
B
Blog RSS Feed
Engineering at Meta
Engineering at Meta
Apple Machine Learning Research
Apple Machine Learning Research
Jina AI
Jina AI
MongoDB | Blog
MongoDB | Blog
腾讯CDC

DEV Community

Authentication Security Deep Dive: From Brute Force to Salted Hashing (With Java Examples) Why AI Systems Don’t Fail — They Drift Spilling beans for how i learn for exam😁"Reinforcement Learning Cheat Sheet" I Replaced Chrome with Safari for AI Browser Automation. Here's What Broke (and What Finally Worked) How Python Borrows Other People's Work The $40 Architecture: Processing 1 Billion API Requests with 99.99% Uptime Vibe Coding: A Workflow Guide (From Zero to SaaS) Most webhook security guides protect the wrong side. The scary part is delivery. Headless CMS for TanStack Start: Build a Blog with Cosmic EU Age Verification App "Hacked in 2 Minutes" — What Actually Happened Comfy Cloud’s delete function does not actually remove files Running AI Models on GPU Cloud Servers: A Beginner Guide Event-driven media intelligence with AWS Step Functions and Bedrock I scored 500 AI prompts across 8 quality dimensions — here's what broke How to Call Google Gemini API from Next.js (Free Tier, No Backend Needed) The Portal Protocol: Reclaiming Human Connection in the Age of AI How to Fix Your Team's Scattered Knowledge Problem With a Self-Hosted Forum Intro to tc Cloud Functors: A Graph-First Mental Model for the Modern Cloud Designing Multi-Tenant Backends With Both Ownership and Team Access I Built a Neumorphic CSS Library with 77+ Components — Here's What I Learned PostgreSQL Performance Optimization: Why Connection Pooling Is Critical at Scale Cómo construí un SaaS multi-rubro para gestionar expensas en Argentina con FastAPI + Vue 3 🚀 I Built an Ethical Hacking Scanner Tool – Open Source Project I Replaced /usage and /context in Claude Code With a Single Statusline A Pythonic Way to Handle Emails (IMAP/SMTP) with Auto-Discovery and AI-Ready Design I Collected 8.9 Million Polymarket Price Points — Here's What I Found About How Markets Really Move EcoTrack AI — Carbon Footprint Tracker & Dashboard Everyone's Using AI. No One Agrees How. 5 self-hosted ebook managers worth trying in 2026 Building Your First AI Agent with LangChain: From Chatbot to Autonomous Assistant
RAG- Understanding of Embedding
Ramya Peruma · 2026-05-18 · via DEV Community

Ramya Perumal

What is Embedding?

After text is split into chunks, the next process is called embedding. In this step, each chunk is converted into vectors (points in vector space). In vector-based RAG systems, chunks are converted into vectors so that semantic search can be performed efficiently.

Why Do We Need to Convert Chunks into Vectors?

The main goal of a RAG application is to achieve semantic search.

Semantic
Example
The word feline is related to the cat family, even though the words are different. Understanding that “feline” and “cat” are related is called semantic understanding.

Similarity
When a user asks a query, semantically related chunks are returned even though the exact words in the chunks may be different.

Semantic Similarity
Semantic similarity combines:

  • Intent
  • Context
  • Meaning

The purpose is to establish relationships between the user query and the documents stored in the RAG system. This allows the system to retrieve relevant information from the database and provide it to the LLM for further processing.

Words that are semantically related are usually stored closer together in multi-dimensional vector space.

Cosine Similarity
To determine how close vectors are to each other, cosine similarity is commonly used.

When a user query arrives:

  1. The query is converted into a vector
  2. Cosine similarity is calculated between the query vector and stored vectors
  3. The closest vectors are retrieved

Retrieval Methodologies

Two major retrieval methodologies are used:

1. KNN (K-Nearest Neighbors)
KNN compares the query vector with all stored vectors one by one to find the nearest neighbors.

Advantage
More accurate retrieval

Disadvantage
Slow for very large datasets

2. ANN (Approximate Nearest Neighbors)
ANN approximately finds the nearest vectors instead of comparing every single point.

This method is mainly used when:

  • The document volume is huge
  • Faster retrieval is required
  • Time constraints exist

ANN improves retrieval speed while sacrificing a small amount of accuracy.

Why Cosine Similarity Instead of Sine or Tangent?

Cosine similarity works effectively because:
If two vectors are very close and highly related, the cosine similarity value approaches 1. If the angle between vectors increases, the cosine similarity value decreases, meaning the vectors are less related

Why Not Sine or Tangent?
For small angles:

  • Sine values remain close to 0
  • Tangent values can fluctuate significantly

These measurements are not stable for semantic comparison. Cosine similarity provides a more reliable way to measure semantic closeness between vectors.

Embedding Dimensions
Embedding models can generate vectors with dimensions ranging from 256 to 3000 or more.

The dimension size depends on the embedding model and the amount of contextual information it captures.

Generally:

  • Higher dimensions capture richer semantic information
  • Lower dimensions are faster and cheaper but may lose context

Types of Embedding Models
Choosing an embedding model completely depends on the application scenario.

1. Based on Query Type

Symmetric Models
Symmetric embedding models are used when the query and the documents are similar in structure and length.

Examples
nomic-embed-text
Qwen embeddings

These are commonly used in semantic search systems.

Asymmetric Models
Asymmetric embedding models are used when:

  • Queries are short
  • Documents are long

Example
Google Gemini embedding models

These models are optimized for retrieving long documents from short queries.

2. Based on Retrieval Type

Dense Embeddings
Dense embeddings mainly focus on semantic meaning.

These embeddings generate dense vectors where most values contain meaningful information.

Examples
Cohere embedding models
ChatGPT OSS 120B embeddings

Advantage
Better semantic understanding

Sparse Embeddings
Sparse embeddings mainly focus on exact keyword matching.

They commonly use the BM25 (Best Match 25) algorithm, which is based on:

  • TF (Term Frequency)
  • IDF (Inverse Document Frequency)

TF-IDF Concepts

TF (Term Frequency)
Measures how many times a word appears in a document.

IDF (Inverse Document Frequency)
Measures how important a word is across the entire document collection.Words that appear too frequently across all documents are considered less important.

Transformer Architecture

The transformer architecture was a major breakthrough for LLMs.
Transformers mainly contain:

  • Encoder
  • Decoder

Encoder
The encoder converts text into embeddings (vectors).

Decoder
The decoder converts embeddings back into human-readable text after processing.

This architecture enables modern LLMs to understand and generate natural language effectively.

Choosing a Vector Database

Chroma
Open source
Easy to set up
Suitable for basic and small-scale applications

FAISS
Better for large document collections
Optimized for high-performance semantic search
Commonly used in production-scale retrieval systems