惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

博客园 - 三生石上(FineUI控件)
G
Google Developers Blog
Vercel News
Vercel News
Cyber Security Advisories - MS-ISAC
Cyber Security Advisories - MS-ISAC
N
Netflix TechBlog - Medium
奇客Solidot–传递最新科技情报
奇客Solidot–传递最新科技情报
钛媒体:引领未来商业与生活新知
钛媒体:引领未来商业与生活新知
Engineering at Meta
Engineering at Meta
B
Blog
博客园_首页
量子位
博客园 - 叶小钗
L
LangChain Blog
T
The Blog of Author Tim Ferriss
云风的 BLOG
云风的 BLOG
Blog — PlanetScale
Blog — PlanetScale
F
Fortinet All Blogs
S
SegmentFault 最新的问题
宝玉的分享
宝玉的分享
D
DataBreaches.Net
雷峰网
雷峰网
The Cloudflare Blog
freeCodeCamp Programming Tutorials: Python, JavaScript, Git & More
Threat Intelligence Blog | Flashpoint
Threat Intelligence Blog | Flashpoint
Last Week in AI
Last Week in AI
P
Proofpoint News Feed
TaoSecurity Blog
TaoSecurity Blog
罗磊的独立博客
MongoDB | Blog
MongoDB | Blog
The GitHub Blog
The GitHub Blog
I
Intezer
H
Help Net Security
The Hacker News
The Hacker News
The Register - Security
The Register - Security
cs.CL updates on arXiv.org
cs.CL updates on arXiv.org
AWS News Blog
AWS News Blog
V
V2EX
Microsoft Security Blog
Microsoft Security Blog
T
Tenable Blog
Spread Privacy
Spread Privacy
A
Arctic Wolf
P
Proofpoint News Feed
T
Threat Research - Cisco Blogs
Schneier on Security
Schneier on Security
C
CERT Recently Published Vulnerability Notes
让小产品的独立变现更简单 - ezindie.com
让小产品的独立变现更简单 - ezindie.com
The Last Watchdog
The Last Watchdog
Latest news
Latest news
T
Troy Hunt's Blog
L
LINUX DO - 热门话题

Databricks

How lakebase architecture delivers 5x faster Postgres writes Why Talent Transformation Is the Missing Focus of Enterprise AI Public Health Intelligence Shouldn't Require a Data Scientist Mean Time to Detect Is a Data Access Problem First-party audience data is the ad sales relationship now Rethinking Distributed Systems for Serverless Performance and Reliability The AI Scaling Gap Hiding in Digital Native Companies 10 trillion samples a day: Scaling beyond traditional monitoring infra at Databricks AI success starts with clean data, not just better models How nOps Rebuilt Their Cloud Optimization Platform on Databricks Lakebase, and Why Other ISVs Should Too Peril Predicts: Precision Payouts for a Volatile World The foundation of AI scalability: one team, one platform, one operating model The Federal Data Paradox: Rich in Data, Poor in Access Driving Budapest Forward: How BKK Uses Databricks to Transform City Mobility LLM Vs AI: A Practical Guide to Differences, Use Cases, and Tools Model Risk Governance Is Not the Same as Risk Intelligence Generative AI for Business: A Complete Strategy and Implementation Guide Data Science vs Data Engineering: Choosing Analysis or Infrastructure AI Applications: Tools, Use Cases, and Platforms MLOps vs DevOps: A Practical Guide for Data Scientists and IT Teams Top Data Warehouse Tools For Modern Data Analytics Unlocking SAP Business Context in Databricks with Semantic Metadata Delta Sharing The marketing activation gap has a fix: Databricks and Stitch partner to turn data infrastructure into marketing performance Alert Fatigue Is a Business Risk Backstage with Lakebase Shipping Faster isn’t Learning Faster Why Your OEE Dashboard Is Lying to You The Turbine That Tried to Tell You It Was Failing Predicting Readmissions Isn't Enough. Acting in Time Is. Clinical Trials Run Longer Than They Have To. That's a Patient Problem Network Quality Is a Revenue Problem, Not a Technical One Shelf Availability Starts with Better Demand Visibility When Predicting the Next Hit Requires More Than Intuition Approximate Answers, Exact Decisions: New Sketch Functions for Analytics Companies Winning with AI Built the Data Layer First Rethinking SQL ETL for modern data platforms Stripe data now available on Databricks via Databricks Marketplace Databricks and Stripe Projects: Infrastructure Built for Agents Agents are ready but your architecture probably isn't Interoperability Between Unity Catalog and Google BigQuery via Catalog Federation Built In, Not Bolted On: What AI-Native Actually Means in Cybersecurity Operationalizing AI for public sector fraud prevention From months to minutes: Building real-time clinical data pipelines with natural language Agentic Data Engineering with Genie Code and Lakeflow Securely send first-party conversion signals with Snapchat Conversions API on Databricks Marketplace How leading tech companies are killing the builder’s tax with Lakebase Inside one of the first production deployments of Lakebase: LangGuard's agentic workflow governance engine The next generation of Databricks Genie Model Risk Management in 2026: A Banker’s Guide to the Revised Interagency Guidance OpenAI GPT-5.5 now available on Databricks, fully-governed through Unity AI Gateway Operational databases: How they work and when to use them Databricks partners with OpenAI on GPT-5.5 Announcing the Public Preview of Lakeflow Designer Are LLM agents good at join order optimization? How conversational analytics removes the BI bottleneck How to transform document activation workflows with Genie and Agent Bricks Beyond the spreadsheet: how Databricks is delivering the modern CFO in Financial Services AI App Development: Guide To Building AI-Powered Apps IoT in Manufacturing: Strategy, Components, Use Cases, and Challenges Stop Hand-Coding Change Data Capture Pipelines Multimodal Data Integration: Production Architectures for Healthcare AI Personalization Strategies for Media Companies A Modern AI Risk Management Framework Introducing the Databricks Excel Add-in for Business Users Real-Time Decisioning for AI Agents: Why you Need a Customer Context Layer First A Practical Guide to LLM Fine Tuning AI Data Transformation Guide for Data Engineers and Data Scientists Concurrency Control in DBMS: How Locking, MVCC and Optimistic Strategies Keep Data Consistent Bridging data science and marketing: Databricks unveils Delta Sharing integration for Adobe Experience Platform and agentic marketing workflows Take Control: Customer-Managed Keys for Lakebase Postgres Get hands on with agents, vibe coding and more at Data+ AI Summit Mercedes-Benz Builds a Cross-Cloud Data Mesh with Delta Sharing and Intelligent Replication, Cutting Costs by 66% What Is a Transactional Database? Introducing Genie Agent Mode Governing coding agent sprawl with Unity AI Gateway Governing Coding Agent Sprawl with Unity AI Gateway Banks Don’t Have an AI Problem – They Have a Data Platform Problem Open Platform, Unified Pipelines: Why dbt on Databricks is Accelerating Why Your Agents Can’t Read Enterprise Documents — and How to Fix It Building with Databricks Document Intelligence and Lakeflow Databricks on Google Cloud: Innovate Faster. Smarter. Together. Introducing the Databricks Connector for Google Sheets: Real-Time, Governed Lakehouse Data in the Sheets Users Love Unity AI Gateway: How to connect agents to external MCPs securely Expanding agent governance with Unity AI Gateway Agentic reasoning in practice: Making sense of structured and unstructured data Agent Bricks: The Governed Enterprise Agent Platform 8 AI and data trends shaping financial services in 2026 Building real-time product search on Databricks Lovable + Databricks: Build Data-Driven Apps at the Speed of Thought Memory scaling for AI agents Powering clinical research innovation: How TriNetX uses Databricks to accelerate drug development Database Branching in Postgres: Git-Style Workflows with Databricks Lakebase How Zalando built a unified data foundation for AI and analytics on Databricks The next era of the open lakehouse: Apache Iceberg™ v3 in Public Preview on Databricks How FSIs eliminate silos between clients, operations, and finance How MakeMyTrip achieved millisecond personalization at scale with Databricks A multi-agent approach to audience intelligence AiChemy: Next-generation agent with MCP, skills and custom data for drug discovery Accelerate business insights with Lakeflow Connect, now with a Free Tier Unlocking Next-Gen Customer Experiences with Data Intelligence for Marketing
What is pgvector?
2026-04-17 · via Databricks

pgvector is an open-source PostgreSQL extension that adds the ability to store, index and search vector embeddings (numerical representations of data). It brings vector data and similarity search into the same system that holds application data, making it possible to power semantic search, recommendations and retrieval-augmented generation (RAG) without relying on an external vector database. pgvector extends Postgres to support these AI-driven use cases.

Many modern AI applications depend on retrieving semantically similar data, not just exact matches. pgvector allows teams to perform this type of retrieval at runtime within their existing Postgres stack. For example, applications often need to retrieve content that is contextually similar to a query, even if the wording is different. This approach is often referred to as cosine similarity, nearest neighbor search or embedding-based search.

This article provides a high-level, educational overview of pgvector rather than detailed implementation guidance.
 

How pgvector works

pgvector adds a new data type to Postgres called vector. It allows embeddings, numerical representations of text, images or other content to be stored alongside relational data without requiring a separate system. These embeddings are typically generated by machine learning models that convert content such as text or images into numerical form.

At a high level, the process is simple. Embeddings are stored in the database. When a query is received, a query embedding is generated from the input, and pgvector returns the records whose vectors are most similar, or closest in meaning to that query. Instead of matching keywords, results are retrieved based on meaning.

pgvector determines similarity using distance metrics:

  • L2 (Euclidean distance): measures the distance between vectors, where smaller values indicate greater similarity
  • Cosine similarity: measures how closely vectors point in the same direction, which often reflects similarity in meaning
  • Inner product: measures alignment between vectors and is often used with normalized embeddings

Key features of pgvector

pgvector includes several features that make vector search practical within Postgres.

  • Indexing: Two index types are supported: HNSW and IVFFlat. HNSW prioritizes query speed and builds a graph structure in memory, but requires more memory. IVFFlat is more memory-efficient and partitions vectors into clusters using a training step, but queries may be slower.
  • Distance metrics: L2, cosine similarity and inner product cover most embedding-based use cases. Hamming distance supports binary vectors, and Jaccard distance supports sparse vectors in more specialized scenarios.
  • Filtered search: Vector similarity can be combined with standard relational filters. For example, results can include the most semantically similar products that are also in stock, within a price range or in a specific category.
  • Hybrid search: pgvector can be paired with Postgres full-text search to blend keyword and semantic search. This allows results to be both contextually relevant and textually precise in a single query.
  • Additional data types:  Options such as halfvec, sparsevec and bit types help reduce memory usage when working with large embedding datasets.

Common use cases for pgvector

pgvector is widely used to power AI-driven application features:

Semantic search and RAG

Applications can retrieve documents or content based on meaning rather than keywords. This is a core component of retrieval-augmented generation (RAG), where large language models use retrieved context to generate accurate, relevant responses. Because pgvector runs similarity search directly within Postgres, this retrieval can happen in real time without requiring a separate system.

Recommendation systems

Items can be matched to past behavior or preferences to support recommendations. This pattern is commonly used for product recommendations, content discovery and personalization in applications. pgvector makes it efficient to identify related items based on patterns in user behavior or content.

Image similarity

Image embeddings can be stored and compared to quickly find visually similar images. This is widely used in media platforms, e-commerce and creative tools. Storing these embeddings alongside application data makes it easier to run similarity searches without additional infrastructure.

Anomaly detection

Outliers can be identified by finding data points that are distant from typical patterns in vector space. This is useful for fraud detection, monitoring and quality control. pgvector enables this by making it easy to compare vectors and detect deviations.

Deduplication

Duplicate or near-duplicate content can be identified, even when it is expressed differently or formatted in different ways. This is important for content management, search quality and data hygiene. Similarity-based comparison makes it possible to detect duplicates beyond exact matches.
 

pgvector vs. dedicated vector databases: when to use each

As vector search becomes part of more applications, teams often face a practical decision: should vector search stay within Postgres, or is a dedicated vector database needed? The answer depends on scale, performance requirements and operational complexity.

The differences can be summarized across key dimensions:

Tool

Operational Complexity

Scalability Ceiling

Hybrid Query Support

Cost

Ecosystem Maturity

pgvector

Lowest (Existing DB)

High (~100M+ vectors)

Best (Native SQL Joins)

Lowest (Included)

High (Postgres ecosystem)

Pinecone

Low (Serverless/SaaS)

Highest (Billions+)

Moderate (Metadata only)

High (Usage-based)

High (AI-specific)

Weaviate

Moderate (Multi-modal)

Very High

High (GraphQL/Vector)

Moderate

High (Open-source)

Qdrant

Moderate (Rust-based)

Very High

High (Filtering-heavy)

Moderate

Growing Fast

pgvector is the natural starting point for teams already using Postgres and operating below the scale ceiling. It works well when vector search is part of an existing application workflow and data volumes or query demands remain manageable. Dedicated vector databases become more relevant when query volume, recall requirements or multi-tenant workloads push beyond what Postgres can efficiently support.

pgvectorscale

pgvectorscale is designed for teams that want to extend how far they can go with pgvector before adopting a dedicated vector database. It addresses the performance and scalability challenges that arise as data volumes and query demands increase, particularly around indexing speed and query latency. By improving how pgvector performs at larger scales, it allows teams to continue using Postgres for longer without re-architecting their systems. This makes it a practical intermediate step for applications approaching the limits of what pgvector can handle on its own.

Limitations and scaling considerations

pgvector is powerful, but it comes with tradeoffs:

  • Performance can degrade at very high vector counts (10M+) without additional optimization or tooling
    • HNSW indexes are memory-intensive, and large deployments may require significant RAM
    • Postgres does not provide built-in sharding for vector workloads, so horizontal scaling requires external tooling or a managed provider
    • Search speed and recall involve a real tradeoff. Recall — the percentage of truly relevant results that are returned — requires deliberate configuration to optimize.

Understanding these limitations helps determine when pgvector is sufficient and when additional infrastructure may be needed.

Getting started with pgvector

pgvector can be installed on macOS and most Linux distributions using standard package managers such as Homebrew. It is also available on many managed Postgres platforms, including AWS RDS, Supabase, Azure Database for PostgreSQL, Google Cloud SQL and Neon.

Installation and setup instructions are available in the official pgvector GitHub repository, which includes step-by-step guidance maintained by the project’s authors.

Databricks customers using Postgres can also reference the Databricks OLTP extensions docs for platform-specific guidance.

pgvector and the modern AI data stack

pgvector operates in the operational serving layer of an AI system, where low-latency retrieval is required at application runtime. It is commonly used to support semantic search, recommendations and retrieval-augmented generation (RAG) within applications.

In contrast, Databricks Mosaic AI Vector Search is better suited for large-scale, batch-processed AI workloads, where data pipelines are managed in the lakehouse. These environments support centralized data processing, large datasets and complex workflows.

These approaches are complementary, and teams often use both across different layers of the stack. pgvector supports real-time application queries, while platforms like Databricks handle large-scale data preparation, embedding generation and model-driven workflows.

Frequently asked questions

Is pgvector a full vector database?
pgvector enables Postgres to store embeddings and perform similarity search directly on that data. However, it is not a purpose-built vector database. Dedicated vector databases provide additional scalability and performance optimizations for larger workloads.

What is the difference between HNSW and IVFFlat in pgvector?
HNSW is optimized for fast query performance and uses an in-memory graph structure, which requires more memory. IVFFlat has a lower memory footprint and organizes vectors into clusters through a training step, but performance can vary depending on the dataset and workload. The choice depends on whether speed or memory efficiency is the priority.

How many vectors can pgvector handle?
pgvector can typically handle millions to tens of millions of vectors, depending on hardware, indexing strategy and query patterns. As datasets grow, performance may decline without careful tuning or additional tooling. Factors such as available memory, index type and query frequency all influence scalability.

Does pgvector support cosine similarity?
Yes, pgvector supports cosine similarity as one of its primary distance metrics. It measures how closely two vectors point in the same direction, which often reflects semantic similarity in embedding-based applications. This makes it well suited for semantic search, recommendation systems and natural language processing.

Is pgvector free and open source?
Yes, pgvector is an open-source project released under a permissive license. It can be used with standard Postgres installations as well as many managed Postgres services. This makes it an accessible starting point for adding vector search capabilities.

Can pgvector do hybrid search?
Yes, pgvector can be combined with Postgres full-text search to support hybrid search. This allows results to balance semantic relevance with keyword matching, improving both accuracy and usability. Hybrid search is especially useful in scenarios such as product search and documentation search, where both meaning and exact terms are important.

Choosing the right vector search approach

pgvector is a practical starting point for any team that wants to add vector search to an existing Postgres application. By storing embeddings alongside relational data and supporting similarity search natively within the database, it removes the operational overhead of managing a separate vector store. For many workloads — semantic search, RAG pipelines, recommendations and anomaly detection — it delivers what teams need without requiring a new system.

As data volumes grow or query demands increase, pgvectorscale can extend how far teams go before a dedicated vector database becomes necessary. For organizations managing large-scale AI workloads across a unified data platform, Databricks Mosaic AI Vector Search offers a complementary approach designed for the lakehouse layer. Together, these tools give teams the flexibility to match their vector search infrastructure to their actual workload requirements — at any scale.