惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

P
Proofpoint News Feed
WordPress大学
WordPress大学
Help Net Security
Help Net Security
Jina AI
Jina AI
Security Latest
Security Latest
Y
Y Combinator Blog
Project Zero
Project Zero
freeCodeCamp Programming Tutorials: Python, JavaScript, Git & More
GbyAI
GbyAI
Know Your Adversary
Know Your Adversary
Threat Intelligence Blog | Flashpoint
Threat Intelligence Blog | Flashpoint
NISL@THU
NISL@THU
Cisco Talos Blog
Cisco Talos Blog
博客园 - 司徒正美
MyScale Blog
MyScale Blog
Cyberwarzone
Cyberwarzone
D
Docker
T
The Blog of Author Tim Ferriss
G
Google Developers Blog
C
CERT Recently Published Vulnerability Notes
B
Blog
L
LangChain Blog
Exploit-DB.com RSS Feed
Exploit-DB.com RSS Feed
SecWiki News
SecWiki News
The Hacker News
The Hacker News
C
Check Point Blog
L
Lohrmann on Cybersecurity
V2EX - 技术
V2EX - 技术
S
Securelist
T
Threat Research - Cisco Blogs
Stack Overflow Blog
Stack Overflow Blog
TaoSecurity Blog
TaoSecurity Blog
云风的 BLOG
云风的 BLOG
Latest news
Latest news
人人都是产品经理
人人都是产品经理
L
LINUX DO - 最新话题
Application and Cybersecurity Blog
Application and Cybersecurity Blog
The Register - Security
The Register - Security
Webroot Blog
Webroot Blog
Simon Willison's Weblog
Simon Willison's Weblog
奇客Solidot–传递最新科技情报
奇客Solidot–传递最新科技情报
Microsoft Security Blog
Microsoft Security Blog
AWS News Blog
AWS News Blog
C
Cybersecurity and Infrastructure Security Agency CISA
cs.AI updates on arXiv.org
cs.AI updates on arXiv.org
小众软件
小众软件
T
Tailwind CSS Blog
cs.CL updates on arXiv.org
cs.CL updates on arXiv.org
宝玉的分享
宝玉的分享
O
OpenAI News

Pinecone

Pinecone Assistant: A Managed Knowledge Layer for Production AI Applications Multi-domain RAG in n8n: why one knowledge base is not enough Allspice Transforms the Culinary Experience with Semantic Search Powered by Pinecone | Pinecone Building RAG workflows in n8n: choosing the right Pinecone node Knowledge needs a meta-knowledge layer Garbage Day: How Pinecone Safely Deletes Billions of Objects at Scale When "Performance" Means Two Different Things Pinecone BYOC: Pinecone in your AWS, GCP, or Azure account, no vendor access True, Relevant, and Wrong: The Applicability Problem in RAG Use the Pinecone Plugin for Claude Code to develop AI Applications Faster Millions at Stake: How Melange's High-Recall Retrieval Prevents Litigation Collapse Powering High-stakes Patent Search at Scale: How Melange Built a Reliable AI System on Pinecone | Pinecone Pinecone Assistant Node in n8n: Turn Any Data Source Into Knowledge RAG with Access Control Pinecone Dedicated Read Nodes are now in Public Preview Inside Pinecone: Slab Architecture New Bulk Data Operations: Update, Delete, and Fetch by Metadata The Hidden Cost of Building: Lessons from Aquant Simplifying Vector Embeddings with Pinecone Integrated Inference Capabilities Pinecone joins Microsoft Marketplace as a Launch Partner GTM Engineering: Clay + Pinecone for AI-powered Sales Outbound Build an AI knowledge assistant with Google Docs and Pinecone Moving Pinecone forward with Ash Ashutosh as CEO and Edo spearheading our growing AI ambitions as Chief Scientist Pinecone Founder Edo Liberty to Spearhead Pinecone’s Growing AI Ambitions; Appoints Ash Ashutosh as CEO to Expand Vector Database Market Leadership Fast, Accurate Retrieval for Creators at Scale: Delphi’s Path Toward a Million Conversational Agents with Pinecone | Pinecone Announcing Pinecone Pioneers: A Program for Builders, Organizers, and Community Leaders What is Context Engineering? Chunking Strategies for LLM Applications Beyond the hype: Why RAG remains essential for modern AI Obviant Makes 30% More Accurate Defense Acquisition Recommendations Combining Sparse and Dense Retrieval with Pinecone | Pinecone Build more knowledgeable AI applications with new LLMs and greater control in Pinecone Assistant #NYTECHWEEK 2025 Retrieval-Augmented Generation (RAG) Accurate and Efficient Metadata Filtering in Pinecone’s Serverless Vector Database | Pinecone Terminal X AI Agents, Powered by Pinecone, Turn Complex Financial Data Into Production-grade Insights at Scale | Pinecone Aquant Delivers Scalable, Expert-level Service Intelligence with Pinecone | Pinecone Cascading retrieval with multi-vector representations: balancing efficiency and effectiveness Vector databases aren't just for large-scale enterprise AI Unveiling DIME: Reproducibility, Scalability, and Formal Analysis of Dimension Importance Estimation for Dense Retrieval | Pinecone Fast and Effective Early Termination for Simple Ranking Functions | Pinecone Domain-specific AI Agents at Scale: CustomGPT.ai Serves 10,000+ Customers with Pinecone | Pinecone Using Pinecone asynchronously with FastAPI A Flexible Resource for Top-Weighted Comparisons Between Sets and Rankings | Pinecone Build secure, scalable agentic AI workflows with Rubrik Annapurna and Pinecone Tool up: Pinecone’s first MCP servers are here Add context to your agent with Pinecone Assistant MCP remote server E2Rank: Efficient and Effective Layer-wise Reranking | Pinecone ColBERT-serve: Efficient Multi-Stage Memory-Mapped Scoring | Pinecone Efficient Constant-Space Multi-Vector Retrieval | Pinecone How Vanguard Worked with Pinecone to Boost Customer Support with Faster Calls and 12% More Accurate Responses | Pinecone Pinecone Named to Fast Company's Annual List of the World's Most Innovative Companies of 2025 Launch Week: Pinecone for agents, search, recommendations, and more Optimizing Pinecone for agents (and more) Retrieval Inference for scale and performance How 1up Turns Sales Reps Into Product Experts with Pinecone | Pinecone Don’t be dense: Launching sparse indexes in Pinecone Unlock High-Precision Keyword Search with pinecone-sparse-english-v0 Evolving Pinecone's architecture to meet the demands of Knowledgeable AI Pinpoint references faster with citation highlights in Pinecone Assistant Bringing the leading vector database to your cloud Getting started with llama-text-embed-v2 Natural Language Counterfactual Explanations for Graphs Using Large Language Models | Pinecone Easily build knowledgeable chat and agent-based applications in minutes with Pinecone Assistant, now generally available How to build an agentic, chat or RAG knowledge system using Pinecone Assistant Real-time RAG with Pinecone and Estuary Flow BigQuery to Pinecone in Real-Time with Estuary Flow Stravito Turns Market and Consumer Data Into Actionable Insights with Pinecone Inference | Pinecone Accelerate prototyping and development with Pinecone Local First-of-its-kind Pinecone Knowledge Platform to Power Best-in-class Retrieval for Customers Introducing integrated inference: Embed, rerank, and retrieve your data with a single API Strengthening security and increasing control with CMEK and API key roles Introducing Pinecone Rerank V0 Introducing cascading retrieval: Unifying dense and sparse with reranking From Idea to Action: How Pinecone Assistant Meaningfully Accelerates AI Business Building AI apps on Azure with Pinecone just got a lot easier Building a reliable, curated, and accurate RAG system with Cleanlab and Pinecone Four features of the Assistant API you aren't using - but should Deploying Pinecone with Infrastructure as Code (IaC) Streamlining CI/CD with Pinecone Local September 2024 Product Update Results of the Big ANN: NeurIPS'23 competition | Pinecone Introducing import from object storage for more efficient data transfer to Pinecone serverless Simplify, enhance, and evaluate RAG development with Pinecone Assistant, now in public preview Vectors and Graphs: Better Together August 2024 Product Update Pinecone Helps Deep Talk Deliver World-Class AI Assistants with Lower Engineering Overhead | Pinecone Assembled Delivers Better, Faster AI- Driven Support with Pinecone | Pinecone Llama 3.1 Agent using LangGraph and Ollama Build knowledgeable AI with Pinecone serverless, now generally available on Microsoft Azure Pinecone serverless is now generally available on Google Cloud, adding knowledge to AI assistants and other applications Accelerating Legal Discovery and Analysis with Pinecone and Voyage AI Bridging Dense and Sparse Maximum Inner Product Search | Pinecone Refine Retrieval Quality with Pinecone Rerank Introducing reranking to Pinecone Inference to simplify building accurate AI July 2024 Product Update Connect to Pinecone within your platform to enable a seamless AI development experience Introducing Pinecone API Versioning RAG Brag with Inkeep Co-Founder Nick Gomez LangGraph and Research Agents Introducing Pinecone Inference to streamline your AI workflow
Full Text Search: Architecture and Design
Amir Ingber, Harry Scholes, Edo Liberty · 2026-05-07 · via Pinecone

Pinecone now supports full text search with a broad set of features: full Lucene-syntax queries, multi-field schemas, BM25 scoring, tokenization in 18 languages, stemming and stop-word removal, phrase matching search, new filters based on text match for vector search, and more.

# Query: keyword search across title/body using Lucene query syntax (BM25 ranking)
index = pc.preview.index(name="transcripts")

response = index.documents.search(
    namespace="animal-kingdom",
    score_by=[
        {
            "type": "query_string",
            "query": 'body:("nice marmot")'
        }
    ],
    top_k=10,
    include_fields=["title"],
)

Why this matters: agents

Agents need to complete tasks, write code, and answer questions. To retrieve the exact data needed, they have to combine different kinds of filters and scoring. Text is a big part of that.

For example, an agent might search through customer support transcriptions for new insights. It looks for user interactions that indicate frustration with wait times — but it has already flagged interactions that contain the words "wait" or "waiting" and wants to exclude those. And, it needs to consider only a specific geographic and time period. Coding agents are great at composing complex queries like this. What they need is an engine that can execute them — one that internally links semantic (vector), text, and metadata filtering in a single API. See below for a detailed API example.

How this enhances Pinecone's search capabilities

Since 2023, Pinecone has supported text search via optimized sparse indexes. Under the hood, sparse indexes encode documents as a list of token IDs and weights — IDs typically come from a tokenizer, weights from a heuristic like BM25. We exposed this as a general mechanism so developers could plug in non-standard tokenizers, train their own lexical models, or use Pinecone's sparse models instead of naive token weight heuristics.

So, sparse indexes let you do more. But they also require you to do more!

Users (and agents) had to deal with tokenization, weighting, maintaining word counts, and ranking heuristics themselves. They needed something that feels more like a search text box and less like a data science project. Full text search is that.

Under the hood: Tantivy

We integrated the Tantivy library for advanced text functionality alongside vectors and metadata. We considered building it in-house and decided against it, for three reasons — ordered by what we think matters most to a developer using this:

  1. Familiar Lucene syntax. Most search experts already know it, and agents know it too. When we ask a coding agent to add a text search clause to a Pinecone query, it usually gets it right on the first try.
  2. 18 languages out of the box. Mature, community-maintained tokenization and stemming for the languages listed below. Getting this right from scratch is hard and never finishes.
  3. Rust and slabs. Tantivy is a Rust library (so are we), and Tantivy segments map cleanly onto Pinecone slabs.

Advanced text search features

Query syntax

Pinecone supports Lucene query syntax via Tantivy's query parser. The most common operators are below; see the docs for the full list.

Feature Syntax Included docs Scoring
Single term fox Docs containing the term BM25 for the term
Phrase "quick brown" Docs containing the exact adjacent phrase BM25 over the phrase as a single term
Boolean AND quick AND dog Docs matching every clause Sum of clause BM25 scores
Boolean OR fox OR dog Docs matching at least one clause Sum of BM25 scores from clauses that matched
Boolean NOT science NOT algorithms Docs matching the positive clause and not the negated one BM25 for the positive clause; the negated clause contributes nothing
Phrase slop "quick fox"~2 Docs where the terms appear within ~N positions of each other BM25 over the phrase; every in-window occurrence counts equally toward TF
Term boost fox^3 Same as the unboosted clause Clause BM25 multiplied by the boost factor

Field scoping is written into the query itself: or . For the full request and response reference, including other ranker types and more query syntax details, see the public preview docs. All control- and data-plane calls require the header while document search is in public preview.

Text match: new filter operations

Pinecone's metadata filter language is already rich — / , / / / , / , , and the boolean combinators / / — and the engine evaluates it efficiently alongside vector search. Full text search adds three new filter operators that work against any text-searchable field: , , and . Each one tokenizes the operand the same way the field is tokenized at index time, then includes or excludes a record based on whether the field matches.

# Each text predicate is a filter clause: it returns true (keep the record)
# or false (drop it). Score comes from score_by, never from the filter.
filter = {"body": {"$match_phrase": "machine learning"}}    # exact adjacent phrase
filter = {"body": {"$match_all":    "machine learning"}}    # all terms, any order
filter = {"body": {"$match_any":    "ml ai neural"}}        # any term

These compose with the existing metadata operators via , , and , so a single block can mix text-match and metadata predicates. The next section pairs that with vector ranking.

Combining vector ranking with a text-match filter

A document either passes the filter or it doesn't; if it passes, its score comes entirely from . The same block works with every ranker — , , (BM25 on a single field), or (multi-field BM25 with Lucene query syntax).

Returning to the agent example from the intro — "find new frustration signals about wait times, but exclude records that already mention or , restrict to EU customers, and only look at recent records" — that becomes a single search:

Public preview is BYO-vector for hybrid: you compute the embedding (or sparse vector) yourself, then pass it in alongside any text-match filters. There is no / integrated-inference field on document-shaped indexes today.

from pinecone import Pinecone
pc = Pinecone(api_key="YOUR_API_KEY")
# Index schema (set when the index was created):
#   from pinecone.preview import SchemaBuilder
#   schema = (
#       SchemaBuilder()
#       .add_dense_vector_field("embedding", dimension=1536, metric="cosine")
#       .add_string_field("body", full_text_search={"language": "en"})
#       .add_string_field("geo", filterable=True)
#       .add_float_field("timestamp", filterable=True)  # epoch seconds
#       .build()
#   )

# 1. Embed the natural-language search intent.
query_embedding = compute_embedding("your wait times are too long")

# 2. Run a single search: the dense vector ranks the surviving documents.
index = pc.preview.index(name="support-conversations")

response = index.documents.search(
    namespace="__default__",
    filter={
        "$and": [
            {"$not": {"body": {"$match_any": "wait waiting"}}},
            {"geo": {"$eq":  "EU"}},
            {"timestamp": {"$gte": 1767225600}}  # 2026-01-01T00:00:00Z
        ]
    },
    
    score_by=[{
        "type": "dense_vector",
        "field": "embedding",
        "values": query_embedding,
    }],
    
    top_k=10,
    include_fields=["body", "geo", "timestamp"],
)

tokenizes its operand the same way the indexed field was tokenized, so the " " filter will match any document containing any of these tokens in the field.

Tokenizer pipeline

Stemming and stop-word filtering are opt-in per text field — both off by default. Configure them on the field's block at index creation (for example, ). When omitted, server defaults apply and the field is tokenized and lowercased only.

Stemming is available for 18 languages: Arabic, Danish, German, Greek, English, Spanish, Finnish, French, Hungarian, Italian, Dutch, Norwegian, Portuguese, Romanian, Russian, Swedish, Tamil, and Turkish.

Stop-word filtering is available for English, Danish, German, Finnish, French, Hungarian, Italian, Dutch, Norwegian, Portuguese, Russian, and Swedish. There is no stop word support for Arabic, Greek, Romanian, Tamil, and Turkish

The analysis chain applied to each field at index and query time, when both filters are enabled, is:

  • SimpleTokenizerRemoveLongFilter(max_term_len)LowerCaserStopWordFilter (when stop_words is enabled)Stemmer (when stemming is enabled)

BM25 scoring

Pinecone uses the common BM25 algorithm for text ranking. For each (query, document) pair, the score is computed from the term frequencies (TF) and inverse document frequencies (IDF) of the terms common to both. TFs are simple counts of the terms in the document. The canonical IDF is global to the searched corpus, given by

where is the total number of documents and is the number of documents containing term .

Ideally, we want to store these document-level statistics so that when we answer queries we simply retrieve the IDF values. We describe below how we make this work with a growing and dynamic dataset in our slab-based architecture.

Design choices

Document ordering

Document order matters a lot in text search. It controls how posting lists compress, how queries traverse the index, and which optimizations are possible. Common choices include time of insertion, cluster or topic, and authority or popularity — each optimizing for a different goal: insertion throughput, index size, or query latency.

Slabs are built offline from immutable batches, so we are free to pick any order. There is no streaming-insertion constraint to preserve, so the reordering costs are low at write time.

When both vectors and text exist per record, we take advantage of the geometric location of vectors in space to inform the ordering of text records. Vectors are grouped by IVF clusters for fast vector search, so nearby vectors belong to semantically similar records. Vector search, metadata filtering, and text search then share the same ordering — scores and predicates from one stage feed the next without translation.

This ordering also compresses posting lists well. When similar documents land on nearby IDs, the gaps between consecutive document IDs in each posting list shrink, and delta-encoded lists become smaller. The effect is well established in the information retrieval literature [1, 2, 3, 4]. Our IVF-based ordering is a cheap approximation of the clustering those papers recommend: we get the compression benefit as a side effect of the ordering we already need for vector search.

BM25 statistics for growing and dynamic data

Exact BM25 needs global term statistics across the whole index. The dataset is growing and changing, though, and Pinecone's slab-based architecture serves queries by partitioning data across slabs and merging slab-level results at query time. To score consistently across slabs, we merge slab-level term statistics for the terms in the query into a single statistics object, then pass that object to every slab as it scores. The merge is fast because it only touches statistics for the query terms.

When all slabs of an index reside on a single query executor, the merged statistics cover the entire index — scores match the canonical global-BM25 ranking. For indexes whose slabs span multiple executors, we avoid the costly scatter-gather across machines and merge statistics only at the executor level. The resulting scores are approximate relative to a single global corpus, but BM25 is logarithmic in term frequencies, so the perturbation barely moves rankings. The tradeoff between local and global statistics in distributed retrieval has been studied since the mid-1990s [5], and per-shard statistics are standard practice in production search engines [6]. Meaningful distortions only show up on very small samples.

Example benchmarks

Dataset. Wikipedia (English, 20231101 dump), every article reduced to a two-field record of title (avg ~3 words) and body (avg ~470 words). We ran the same query suite against five corpus sizes — 1k, 10k, 100k, 1M, and the full 6.4M articles.

Queries.

  • BM25 set: 5,000 Lucene-syntax queries spanning 20 categories — single terms, phrases, slop, term boosts, cross-field AND/OR, complex booleans combining 4–8 clauses. Generated to mirror the long, multi-clause queries that RAG and agent workloads produce.
  • Filtered set: 500 BM25 queries each paired with a text-match filter ($match_phrase / $match_all / $match_any over title and/or body, some combined with metadata predicates). Filter selectivity ranges from ~1 doc to ~60% of the corpus, exercising both the highly-selective and the broadly-permissive ends of the curve.

Setup. A single Pinecone index per corpus size, b1 dedicated node type, 1 shard / 1 replica. Latency is measured client side. Client and index are in the same region (AWS us-east-1). Top-k = 100. Queries are sent sequentially, one at a time. Latency reflects end-to-end per-query latency, not batching throughput.

SizeBM25 p50 (ms)BM25 p99Filter p50 (ms)Filter p99
1k6.09.66.19.6
10k6.09.76.252.3
100k6.612.26.650.3
1M9.128.39.460.0
6.4M22.784.015.8136.0

As a benchmark methodology — not a public-API guarantee — we measure accuracy against an external brute-force baseline that scores every (query, doc) pair with the canonical (single-machine, global) BM25 formula and ranks them exactly. Recall is the standard at top-k = 100.

Across our benchmark corpora, both plain BM25 and BM25-with-text-match-filter queries deliver mean recall well above 95% on indexes from 1k up to millions of documents. The small gap is concentrated in queries where many documents share an identical BM25 score at the top-k boundary. In such cases, the identity of the documents ending in the top-k depends on the scan order between slabs that can be arbitrary, and potentially different than the brute-force scan order.

Document search is in public preview. The Python SDK exposes it under , and every request must include the header .

pip install --upgrade pinecone
from pinecone import Pinecone
from pinecone.preview import SchemaBuilder

pc = Pinecone(api_key="YOUR_API_KEY")

# 1. Define the index schema: two FTS string fields.
schema = (
    SchemaBuilder()
    .add_string_field("title", full_text_search={"language": "en"})
    .add_string_field("body",  full_text_search={"language": "en"})
    .build()
)

# 2. Create a serverless preview index. Wait for status.ready before continuing.
pc.preview.indexes.create(
    name="quickstart",
    schema=schema,
    read_capacity={"mode": "OnDemand"},
)

# 3. Upsert documents to the default namespace.
index = pc.preview.index(name="quickstart")
index.documents.batch_upsert(
    namespace="__default__",
    documents=[
        {"_id": "1", "title": "The Big Lebowski", "body": "nice marmot"},
        {"_id": "2", "title": "Fargo",            "body": "wood chipper"},
    ],
)

# 4. Search across both fields with Lucene query syntax.
response = index.documents.search(
    namespace="__default__",
    top_k=10,
    score_by=[{"type": "query_string", "query": 'body:("nice marmot")'}],
    include_fields=["title"],
)

For the full reference, see the docs.

References

  1. Blandford and Blelloch, Index Compression through Document Reordering, DCC 2002.
  2. Silvestri, Sorting Out the Document Identifier Assignment Problem, ECIR 2007.
  3. Yan, Ding, and Suel, Inverted Index Compression and Query Processing with Optimized Document Ordering, WWW 2009.
  4. Dhulipala et al., Compressing Graphs and Indexes with Recursive Graph Bisection, KDD 2016.
  5. Viles and French, Dissemination of Collection Wide Information in a Distributed Information Retrieval System, SIGIR 1995.
  6. Elastic, Practical BM25 — Part 1: How Shards Affect Relevance Scoring in Elasticsearch.