惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

雷峰网
雷峰网
D
Darknet – Hacking Tools, Hacker News & Cyber Security
P
Proofpoint News Feed
Spread Privacy
Spread Privacy
C
Cisco Blogs
L
Lohrmann on Cybersecurity
宝玉的分享
宝玉的分享
I
Intezer
aimingoo的专栏
aimingoo的专栏
Cisco Talos Blog
Cisco Talos Blog
The Register - Security
The Register - Security
GbyAI
GbyAI
C
CERT Recently Published Vulnerability Notes
Apple Machine Learning Research
Apple Machine Learning Research
U
Unit 42
Cyberwarzone
Cyberwarzone
爱范儿
爱范儿
I
InfoQ
博客园_首页
让小产品的独立变现更简单 - ezindie.com
让小产品的独立变现更简单 - ezindie.com
量子位
P
Palo Alto Networks Blog
Microsoft Azure Blog
Microsoft Azure Blog
H
Hackread – Cybersecurity News, Data Breaches, AI and More
有赞技术团队
有赞技术团队
T
Tailwind CSS Blog
腾讯CDC
阮一峰的网络日志
阮一峰的网络日志
NISL@THU
NISL@THU
T
Threatpost
T
The Blog of Author Tim Ferriss
云风的 BLOG
云风的 BLOG
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻
S
Schneier on Security
Security Latest
Security Latest
Martin Fowler
Martin Fowler
H
Help Net Security
小众软件
小众软件
The Hacker News
The Hacker News
Know Your Adversary
Know Your Adversary
博客园 - 司徒正美
人人都是产品经理
人人都是产品经理
T
The Exploit Database - CXSecurity.com
Jina AI
Jina AI
Engineering at Meta
Engineering at Meta
The GitHub Blog
The GitHub Blog
P
Proofpoint News Feed
IT之家
IT之家
WordPress大学
WordPress大学
S
Securelist

DEV Community

Authentication Security Deep Dive: From Brute Force to Salted Hashing (With Java Examples) Why AI Systems Don’t Fail — They Drift Spilling beans for how i learn for exam😁"Reinforcement Learning Cheat Sheet" I Replaced Chrome with Safari for AI Browser Automation. Here's What Broke (and What Finally Worked) How Python Borrows Other People's Work The $40 Architecture: Processing 1 Billion API Requests with 99.99% Uptime Vibe Coding: A Workflow Guide (From Zero to SaaS) Most webhook security guides protect the wrong side. The scary part is delivery. Headless CMS for TanStack Start: Build a Blog with Cosmic EU Age Verification App "Hacked in 2 Minutes" — What Actually Happened Comfy Cloud’s delete function does not actually remove files Running AI Models on GPU Cloud Servers: A Beginner Guide Event-driven media intelligence with AWS Step Functions and Bedrock I scored 500 AI prompts across 8 quality dimensions — here's what broke How to Call Google Gemini API from Next.js (Free Tier, No Backend Needed) The Portal Protocol: Reclaiming Human Connection in the Age of AI How to Fix Your Team's Scattered Knowledge Problem With a Self-Hosted Forum Intro to tc Cloud Functors: A Graph-First Mental Model for the Modern Cloud Designing Multi-Tenant Backends With Both Ownership and Team Access I Built a Neumorphic CSS Library with 77+ Components — Here's What I Learned PostgreSQL Performance Optimization: Why Connection Pooling Is Critical at Scale Cómo construí un SaaS multi-rubro para gestionar expensas en Argentina con FastAPI + Vue 3 🚀 I Built an Ethical Hacking Scanner Tool – Open Source Project I Replaced /usage and /context in Claude Code With a Single Statusline A Pythonic Way to Handle Emails (IMAP/SMTP) with Auto-Discovery and AI-Ready Design I Collected 8.9 Million Polymarket Price Points — Here's What I Found About How Markets Really Move EcoTrack AI — Carbon Footprint Tracker & Dashboard Everyone's Using AI. No One Agrees How. 5 self-hosted ebook managers worth trying in 2026 Building Your First AI Agent with LangChain: From Chatbot to Autonomous Assistant Common SOC 2 Failures (Real World) Stop Vibe-Checking Your AI App: A Practical Guide to Evals How to Use SonarQube and SonarScanner Locally to Level Up Your Code Quality Your Next To-Do App Is Dead — I Replaced Mine with an OpenClaw AI Sign a Nostr event in 60 lines of Python using coincurve — no nostr-sdk, no nbxplorer, no rust toolchain ITGC Audit Explained Like You’re in Big 4 Patch Tuesday abril 2026: Microsoft parcha 163 vulnerabilidades y un zero-day en SharePoint Stop scraping everything: a better way to track competitor price changes Listing on MCPize + the Official MCP Registry while routing payments OUTSIDE the marketplace — how I kept 100% of my x402 revenue Building an AI-Powered Risk Intelligence System Using Serverless Architecture Why We Ripped Function Overloading Out of Our AI Toolchain Testing AI-Generated Code: How to Actually Know If It Works SaaS Churn Is Killing Your Business. Here Is What to Do About It (Without a Support Team) The Speed of AI Is No Longer Linear - And Self-Improving Models Are Why How to Implement RBAC for MCP Tools: A Practical Guide for Engineering Teams From Standard Quote to Persuasive Proposal: AI Automation for Arborists I built a CLI that scaffolds complete multi-tenant SaaS apps Axios CVE-2025–62718: The Silent SSRF Bug That Could Be Hiding in Your Node.js App Right Now The dashboard that ended our friendship Data Pipelines Explained Simply (and How to Build Them with Python) The Hidden Cost of AI Systems Nobody Talks About. undefined vs undeclared, and how typeof behaves Switching from file-based jobs to NATS/Kafka in Rust without changing code io_uring Adventures: Rust Servers That Love Syscalls Why Agentic AI is Killing the Traditional Database The POUR principles of web accessibility for developers and designers Quantum Neural Network 3D — A Deep Dive into Interactive WebGL Visualization How To Install Caveman In Codex On macOS And Windows Automation Pipeline Reliability: Why Your Workflow Breaks When Nobody Is Watching I Built an 'Open World' AI Coding Agent — It Works From ANY Folder From Freelancing to Product: A Tech Service Company's SaaS Transformation China's AI Giants: Adding Tencent Hunyuan & ByteDance Doubao to AI University (74 Providers) On the Vibe Coders and Their Lies clerk: Auto-Summarize Your Claude Code Sessions AI Weekly — 2026/04/10–04/17 | The Model Lockdown Is Here, but the Toolchain Is the Real Battleground AI 週報 — 2026/04/10–2026/04/17 模型封鎖潮來了,但工具鏈才是真戰場 Maybe this is how Open-Source apps are born... 🚀 Fine-Tune LLMs with LoRA and QLoRA: 2026 Guide tRPC v11 + Next.js App Router: End-to-End Type Safety Without the Boilerplate ShadCN UI in 2026: Why I Stopped Installing Component Libraries and Started Owning My Components SaaS Billing in React Server Components: Stripe + Supabase Without a Single `useEffect` Join our DEV Weekend Challenge — $1,000 in Prizes Across TEN winners! Submissions Due April 20 at 6:59 AM UTC. Implementing FSRS Spaced Repetition in Flutter + Supabase — Adding Memory Science to an AI Learning App "I Texted My Localhost From the Train — Claude Code Fixed the Bug Before I Got Home" I Built a Sales Prep AI and It Went Deeper Than Expected Design to Code #2: One JSON, Eleven Outputs Solving the 100M-Row Problem: A Summary Table Pattern for High-Volume Push Notification Logs Flutter Web With Wasm: What Actually Changes For Developers I Built 50 Royalty-Free Soundtracks for My Side Project in a Weekend Using AI Music Generation The Vibe Coding Security Checklist: 7 Things to Check Before You Ship Stop Letting Googlebot Guess Fix Your React App's SEO Right Desconstruindo o Streaming do LinkedIn: Como Criar um Engine de Extração de Vídeo de Alta Performance com HLS e FFmpeg (EDA Part-1) EDA (Exploratory Data Analysis) Explained With Real Life — Why Looking at Your Data Is the Most Important Step in Machine Learning Brand Relationship Management at Scale: Our 4-Touch Outreach System for 200+ Brands Why String.fromEnvironment() Might Return an Empty String in Dart JGuardrails 1.0.0 — Hardening Java LLM Apps Against Jailbreaks, Toxicity, and Prompt Injection Plan and Schedule a Full Week of Threads Content From One Claude Conversation Coding Cat Oran Ep3, Five Tables Changed Everything Updated: BFF Pattern I'm done watching freelancers get buried by 200 proposals. So I'm building the alternative. This is my first post BFS Algorithm in Java Step by Step Tutorial with Examples Tracking LLM Pricing Monthly: An Open Dataset for 22 AI Models How We Measure Content ROI on a Comparison Site: Revenue Attribution Without Perfect Data Introducing Nova AI Ops: The AI-Native Operating System for SRE Teams I built a free desktop video downloader for Windows — Grabbit How Talkie OCR Helps Vision-Impaired & Dyslexic Users Read the World Around Them VRCFaceTracking安装和iPhone面捕配置教程,有bug Even CrowdStrike Can't See Your Agents The Automation Gold Rush: What n8n Workflows and Claude Are Opening Up for Developers Right Now
RAG Series (6): Vector Databases — Storage and Retrieval Infrastructure
WonderLab · 2026-05-04 · via DEV Community

Why Do We Need Specialized Vector Databases?

In the first five articles, we figured out how to chunk documents and generate embeddings. Now where do these vectors live, and how are they efficiently retrieved?

You might wonder: "Can't I just store vectors in Redis or PostgreSQL?"

No — traditional databases are designed for exact queries (e.g., WHERE id = 123), while vector retrieval is Approximate Nearest Neighbor (ANN) search: given a query vector, quickly find the Top-K most similar vectors among hundreds of millions of document vectors. Traditional database indexes (B-trees, hash tables) are powerless against this type of "similarity query."

Example:

  • Traditional query: Find user with id=42 → O(1) or O(log n)
  • Vector query: Find 10 people most similar to user A → requires comparing against all vectors, brute-force O(n) is too slow

Vector databases use specialized ANN indexes (HNSW, IVF, etc.) to reduce O(n) to O(log n), completing similarity searches across billions of vectors in milliseconds.


Three Core Capabilities of Vector Databases

1. Vector Storage

Store massive numbers of vectors (millions to billions), with each vector attached to raw text and metadata. Support incremental writes and deletions.

2. ANN Approximate Nearest Neighbor Search

Core algorithms:

Algorithm Principle Pros Cons
HNSW Hierarchical Navigable Small World graph, multi-layer structure, search from coarse to fine Fast, accurate, supports dynamic updates Higher memory usage
IVF Inverted File, partition vector space into clusters, find nearest cluster first then search Memory-efficient, good for static data Adding/removing vectors requires index rebuild
Flat Brute-force full comparison 100% accurate Extremely slow, only for small datasets

Recommendation: Use Flat for development (simplicity), HNSW for production (best performance).

3. Metadata Filtering

This is the key capability that distinguishes vector databases from pure vector libraries (like FAISS). You can do two things simultaneously:

  • Find semantically relevant content using vector similarity
  • Apply exact filtering with metadata conditions (e.g., time > 2024-01-01 AND category = "technical")
# Example: Only retrieve technical documents after 2024
results = vectorstore.similarity_search(
    query="microservices monitoring",
    k=5,
    filter={
        "category": "technical",
        "year": {"$gte": 2024}
    }
)

Enter fullscreen mode Exit fullscreen mode


Mainstream Vector Database Comparison

Five Databases at a Glance

Database Positioning Deployment Index Metadata Filter Best For
Chroma Dev/Prototyping Local/Embedded HNSW Local quick validation, small projects
Qdrant Production self-hosted Docker/K8s HNSW Enterprise choice, strong performance & filtering
Weaviate Hybrid search Docker/Managed HNSW When BM25 + vector hybrid retrieval is needed
pgvector PG extension PostgreSQL plugin HNSW/IVF Existing PG environment, avoid new databases
Pinecone Managed cloud Fully managed Auto Zero ops, quick launch

Detailed Analysis

Chroma — The Developer's Best Friend

from langchain_chroma import Chroma

# Embedded operation, zero configuration
vectorstore = Chroma.from_documents(
    documents=chunks,
    embedding=embeddings,
    persist_directory="./chroma_db"
)

Enter fullscreen mode Exit fullscreen mode

  • ✅ Zero config, works after pip install
  • ✅ Supports persistence to local disk
  • ❌ Single-machine performance limited, not for high concurrency
  • ❌ Weak distributed capabilities

Qdrant — The Enterprise Choice for Production

from langchain_qdrant import Qdrant
from qdrant_client import QdrantClient

client = QdrantClient(url="http://localhost:6333")
vectorstore = Qdrant(
    client=client,
    collection_name="docs",
    embeddings=embeddings,
)

Enter fullscreen mode Exit fullscreen mode

  • ✅ Written in Rust, extremely high performance
  • ✅ Very powerful metadata filtering expressions
  • ✅ Supports distributed clustering
  • ✅ Cloud-hosted version available
  • ❌ Requires additional service deployment

Weaviate — The Hybrid Search Specialist

from langchain_weaviate import WeaviateVectorStore
import weaviate

client = weaviate.connect_to_local()
vectorstore = WeaviateVectorStore(
    client=client,
    index_name="Docs",
    text_key="text",
    embedding=embeddings,
)

Enter fullscreen mode Exit fullscreen mode

  • ✅ Native BM25 + vector hybrid retrieval support
  • ✅ Built-in vectorization module (optional)
  • ❌ Higher resource consumption
  • ❌ Steep learning curve

pgvector — A Blessing for PostgreSQL Users

-- Install extension in PostgreSQL
CREATE EXTENSION vector;

-- Create table with vector column
CREATE TABLE documents (
    id SERIAL PRIMARY KEY,
    content TEXT,
    embedding vector(1024),
    category VARCHAR(50)
);

-- Create HNSW index
CREATE INDEX ON documents USING hnsw (embedding vector_cosine_ops);

Enter fullscreen mode Exit fullscreen mode

from langchain_community.vectorstores import PGVector

vectorstore = PGVector(
    connection_string="postgresql://user:pass@localhost/db",
    embedding_function=embeddings,
    collection_name="docs",
)

Enter fullscreen mode Exit fullscreen mode

  • ✅ Seamless integration with existing PG databases
  • ✅ Full SQL expressiveness
  • ✅ Transaction support (ACID)
  • ❌ Vector retrieval performance lower than dedicated databases
  • ❌ PostgreSQL itself becomes bottleneck at large scale

Pinecone — For Those Who Want Zero Ops

from langchain_pinecone import PineconeVectorStore
from pinecone import Pinecone

pc = Pinecone(api_key="your-key")
index = pc.Index("docs")
vectorstore = PineconeVectorStore(index=index, embedding=embeddings)

Enter fullscreen mode Exit fullscreen mode

  • ✅ Fully managed, zero operations
  • ✅ Auto-scaling
  • ✅ Good metadata filtering support
  • ❌ Higher price
  • ❌ Data lock-in (high migration cost)

Practical: Chroma (Development) vs Qdrant (Production)

Scenario

Let's build RAG for a technical blog system:

  • Development: Use Chroma for quick validation, local execution
  • Production: Migrate to Qdrant, supporting multi-tenancy and metadata filtering

Development Phase — Chroma

from langchain_chroma import Chroma
from langchain_openai import OpenAIEmbeddings

embeddings = OpenAIEmbeddings(
    model="BAAI/bge-large-zh-v1.5",
    api_key=os.getenv("SILICONFLOW_API_KEY"),
    base_url="https://api.siliconflow.cn/v1",
    chunk_size=32,
)

# Create local vector store
vectorstore = Chroma.from_documents(
    documents=chunks,
    embedding=embeddings,
    persist_directory="./chroma_db",
    collection_metadata={"hnsw:space": "cosine"}
)

# Search
results = vectorstore.similarity_search("microservices decomposition principles", k=3)
for doc in results:
    print(f"Source: {doc.metadata['source']}")
    print(f"Content: {doc.page_content[:100]}...")
    print()

# Persistence (Chroma auto-saves)
# Next time load with:
# vectorstore = Chroma(persist_directory="./chroma_db", embedding_function=embeddings)

Enter fullscreen mode Exit fullscreen mode

Production Phase — Qdrant

from langchain_qdrant import Qdrant
from qdrant_client import QdrantClient, models

# Connect to Qdrant service
client = QdrantClient(url="http://localhost:6333")

# Create Collection (like a database table)
client.create_collection(
    collection_name="blog_docs",
    vectors_config=models.VectorParams(
        size=1024,  # BGE-large-zh dimensions
        distance=models.Distance.COSINE,
    ),
)

# Write data
vectorstore = Qdrant(
    client=client,
    collection_name="blog_docs",
    embeddings=embeddings,
)

vectorstore.add_documents(documents=chunks)

# Search with metadata filtering
results = vectorstore.similarity_search(
    query="microservices monitoring",
    k=5,
    filter=models.Filter(
        must=[
            models.FieldCondition(
                key="category",
                match=models.MatchValue(value="microservices")
            ),
            models.FieldCondition(
                key="year",
                range=models.Range(gte=2024)
            ),
        ]
    )
)

Enter fullscreen mode Exit fullscreen mode

Migrating from Chroma to Qdrant

# 1. Export from Chroma
chroma_store = Chroma(
    persist_directory="./chroma_db",
    embedding_function=embeddings
)
all_docs = chroma_store.get()

# 2. Import to Qdrant
qdrant_store = Qdrant(
    client=qdrant_client,
    collection_name="blog_docs",
    embeddings=embeddings,
)

# 3. Batch write (Qdrant supports efficient batch import)
from langchain_core.documents import Document

docs = [
    Document(page_content=text, metadata=meta)
    for text, meta in zip(all_docs["documents"], all_docs["metadatas"])
]
qdrant_store.add_documents(docs)

Enter fullscreen mode Exit fullscreen mode


How to Choose a Similarity Algorithm?

When comparing two vectors, the database needs a "distance metric." Three common ones:

Cosine Similarity

cosine(A, B) = (A · B) / (||A|| × ||B||)

Enter fullscreen mode Exit fullscreen mode

  • Measures: Cosine of the angle between two vectors
  • Characteristics: Only cares about direction, not magnitude
  • Best for: Text semantic similarity (most common choice)
  • Range: -1 (opposite) to 1 (identical), typically > 0.7 is similar

Dot Product

dot(A, B) = A · B = Σ(Ai × Bi)

Enter fullscreen mode Exit fullscreen mode

  • Measures: Sum of element-wise products
  • Characteristics: Considers both direction and magnitude
  • Best for: Recommendation systems (preference intensity matters)
  • Note: If vectors aren't normalized, dot product is affected by vector length

Euclidean Distance

euclidean(A, B) = √Σ(Ai - Bi)²

Enter fullscreen mode Exit fullscreen mode

  • Measures: Straight-line distance between two points
  • Characteristics: Absolute distance, sensitive to numerical differences
  • Best for: Image retrieval, numerical feature scenarios
  • Range: 0 (identical) to ∞, smaller is more similar

Selection Guide

Scenario Recommended Algorithm Reasoning
Text semantic retrieval Cosine Similarity Standard choice, insensitive to vector length
Recommendation systems Dot Product Considers user interest intensity
Image retrieval Euclidean Distance Pixel-level differences are more intuitive
Uncertain Cosine Similarity Safest choice

⚠️ Important: The embedding model and similarity algorithm must match! BGE models recommend cosine similarity, and OpenAI text-embedding-3 series also recommend cosine similarity.


Metadata Filtering: From "Needle in a Haystack" to "Precise Targeting"

Why Metadata Filtering?

Suppose your knowledge base has 100,000 documents covering tech, product, operations, and sales. A user asks: "What are this year's sales targets?"

Pure vector retrieval might return:

  • ✅ Sales department 2024 target document
  • ❌ Operations document mentioning "sales" workflows
  • ❌ Technical document about "sales system architecture"

Adding metadata filter {"department": "sales", "year": 2024} precisely scopes the search.

Metadata Filtering in LangChain

from langchain_chroma import Chroma

# Write with metadata
docs = [
    Document(
        page_content="2024 sales target: 30% revenue growth...",
        metadata={"department": "sales", "year": 2024, "type": "target"}
    ),
    Document(
        page_content="Sales system uses Redis caching...",
        metadata={"department": "tech", "year": 2024, "type": "architecture"}
    ),
]

vectorstore = Chroma.from_documents(docs, embeddings)

# Search with filter
results = vectorstore.similarity_search(
    "sales targets",
    k=3,
    filter={"department": "sales", "year": 2024}
)

Enter fullscreen mode Exit fullscreen mode

Qdrant Advanced Filtering Expressions

from qdrant_client import models

filter = models.Filter(
    must=[  # AND conditions
        models.FieldCondition(key="department", match=models.MatchValue(value="sales")),
        models.FieldCondition(key="year", range=models.Range(gte=2024)),
    ],
    should=[  # OR conditions
        models.FieldCondition(key="type", match=models.MatchValue(value="target")),
        models.FieldCondition(key="type", match=models.MatchValue(value="summary")),
    ],
    must_not=[  # NOT conditions
        models.FieldCondition(key="status", match=models.MatchValue(value="draft")),
    ]
)

Enter fullscreen mode Exit fullscreen mode


Selection Summary

By Scenario

Scenario Recommended Database Reasoning
Local dev / quick prototype Chroma Zero config, pip install and go
Production self-hosted Qdrant Best performance, most flexible filtering, stable Rust
Need hybrid search (BM25 + vector) Weaviate Native support for both retrieval types
Existing PostgreSQL pgvector No new components, full SQL expressiveness
Zero ops desired Pinecone Fully managed, auto-scales
Ultra-large scale (billion vectors) Milvus Designed for massive vector scale (not covered in detail)

Migration Path from Dev to Production

Phase 1: Development Validation
    └── Chroma (local embedded, zero config)
            ↓
Phase 2: Testing Environment
    └── Qdrant Docker (single node, validate functionality)
            ↓
Phase 3: Production Launch
    └── Qdrant Cluster / Pinecone Managed (high availability)

Enter fullscreen mode Exit fullscreen mode


Summary

This article covered the core knowledge of vector databases:

  1. Why vector databases are needed — ANN retrieval is something traditional databases can't do
  2. Three core capabilities — storage, ANN retrieval, metadata filtering
  3. Five database comparison — Chroma, Qdrant, Weaviate, pgvector, Pinecone
  4. Practical code — complete examples for Chroma development and Qdrant production
  5. Similarity algorithms — choosing among cosine, dot product, and Euclidean distance
  6. Metadata filtering — from "needle in a haystack" to "precise targeting"

Key Insight: The best vector database isn't the most expensive one — it's the one that fits your scenario. Use Chroma for development, Qdrant for production, pgvector if you have PostgreSQL, Pinecone if you want zero ops. No silver bullet, only the right fit.


References