惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

IT之家
IT之家
腾讯CDC
让小产品的独立变现更简单 - ezindie.com
让小产品的独立变现更简单 - ezindie.com
K
Kaspersky official blog
V
Visual Studio Blog
博客园 - 聂微东
Recent Commits to openclaw:main
Recent Commits to openclaw:main
cs.AI updates on arXiv.org
cs.AI updates on arXiv.org
云风的 BLOG
云风的 BLOG
T
Tailwind CSS Blog
C
Check Point Blog
H
Heimdal Security Blog
The GitHub Blog
The GitHub Blog
Google Online Security Blog
Google Online Security Blog
P
Proofpoint News Feed
AI
AI
The Register - Security
The Register - Security
SecWiki News
SecWiki News
Help Net Security
Help Net Security
T
Troy Hunt's Blog
V
V2EX
T
Tenable Blog
Threat Intelligence Blog | Flashpoint
Threat Intelligence Blog | Flashpoint
Forbes - Security
Forbes - Security
P
Privacy International News Feed
Microsoft Azure Blog
Microsoft Azure Blog
A
About on SuperTechFans
Recorded Future
Recorded Future
C
Cybersecurity and Infrastructure Security Agency CISA
博客园 - 司徒正美
博客园 - 叶小钗
Y
Y Combinator Blog
人人都是产品经理
人人都是产品经理
S
Security @ Cisco Blogs
罗磊的独立博客
Apple Machine Learning Research
Apple Machine Learning Research
L
LINUX DO - 最新话题
V2EX - 技术
V2EX - 技术
The Cloudflare Blog
Jina AI
Jina AI
T
The Exploit Database - CXSecurity.com
L
Lohrmann on Cybersecurity
Webroot Blog
Webroot Blog
美团技术团队
N
News and Events Feed by Topic
小众软件
小众软件
Google DeepMind News
Google DeepMind News
G
GRAHAM CLULEY
阮一峰的网络日志
阮一峰的网络日志
B
Blog

Pinecone

Pinecone Assistant: A Managed Knowledge Layer for Production AI Applications Multi-domain RAG in n8n: why one knowledge base is not enough Allspice Transforms the Culinary Experience with Semantic Search Powered by Pinecone | Pinecone Building RAG workflows in n8n: choosing the right Pinecone node Knowledge needs a meta-knowledge layer Garbage Day: How Pinecone Safely Deletes Billions of Objects at Scale When "Performance" Means Two Different Things Pinecone BYOC: Pinecone in your AWS, GCP, or Azure account, no vendor access True, Relevant, and Wrong: The Applicability Problem in RAG Use the Pinecone Plugin for Claude Code to develop AI Applications Faster Millions at Stake: How Melange's High-Recall Retrieval Prevents Litigation Collapse Powering High-stakes Patent Search at Scale: How Melange Built a Reliable AI System on Pinecone | Pinecone Pinecone Assistant Node in n8n: Turn Any Data Source Into Knowledge RAG with Access Control Pinecone Dedicated Read Nodes are now in Public Preview Inside Pinecone: Slab Architecture New Bulk Data Operations: Update, Delete, and Fetch by Metadata The Hidden Cost of Building: Lessons from Aquant Simplifying Vector Embeddings with Pinecone Integrated Inference Capabilities Pinecone joins Microsoft Marketplace as a Launch Partner GTM Engineering: Clay + Pinecone for AI-powered Sales Outbound Build an AI knowledge assistant with Google Docs and Pinecone Moving Pinecone forward with Ash Ashutosh as CEO and Edo spearheading our growing AI ambitions as Chief Scientist Pinecone Founder Edo Liberty to Spearhead Pinecone’s Growing AI Ambitions; Appoints Ash Ashutosh as CEO to Expand Vector Database Market Leadership Fast, Accurate Retrieval for Creators at Scale: Delphi’s Path Toward a Million Conversational Agents with Pinecone | Pinecone Announcing Pinecone Pioneers: A Program for Builders, Organizers, and Community Leaders What is Context Engineering? Chunking Strategies for LLM Applications Beyond the hype: Why RAG remains essential for modern AI Obviant Makes 30% More Accurate Defense Acquisition Recommendations Combining Sparse and Dense Retrieval with Pinecone | Pinecone Build more knowledgeable AI applications with new LLMs and greater control in Pinecone Assistant #NYTECHWEEK 2025 Retrieval-Augmented Generation (RAG) Accurate and Efficient Metadata Filtering in Pinecone’s Serverless Vector Database | Pinecone Terminal X AI Agents, Powered by Pinecone, Turn Complex Financial Data Into Production-grade Insights at Scale | Pinecone Aquant Delivers Scalable, Expert-level Service Intelligence with Pinecone | Pinecone Cascading retrieval with multi-vector representations: balancing efficiency and effectiveness Vector databases aren't just for large-scale enterprise AI Unveiling DIME: Reproducibility, Scalability, and Formal Analysis of Dimension Importance Estimation for Dense Retrieval | Pinecone Fast and Effective Early Termination for Simple Ranking Functions | Pinecone Domain-specific AI Agents at Scale: CustomGPT.ai Serves 10,000+ Customers with Pinecone | Pinecone Using Pinecone asynchronously with FastAPI A Flexible Resource for Top-Weighted Comparisons Between Sets and Rankings | Pinecone Build secure, scalable agentic AI workflows with Rubrik Annapurna and Pinecone Tool up: Pinecone’s first MCP servers are here Add context to your agent with Pinecone Assistant MCP remote server E2Rank: Efficient and Effective Layer-wise Reranking | Pinecone ColBERT-serve: Efficient Multi-Stage Memory-Mapped Scoring | Pinecone Efficient Constant-Space Multi-Vector Retrieval | Pinecone How Vanguard Worked with Pinecone to Boost Customer Support with Faster Calls and 12% More Accurate Responses | Pinecone Pinecone Named to Fast Company's Annual List of the World's Most Innovative Companies of 2025 Launch Week: Pinecone for agents, search, recommendations, and more Optimizing Pinecone for agents (and more) Retrieval Inference for scale and performance How 1up Turns Sales Reps Into Product Experts with Pinecone | Pinecone Don’t be dense: Launching sparse indexes in Pinecone Unlock High-Precision Keyword Search with pinecone-sparse-english-v0 Pinpoint references faster with citation highlights in Pinecone Assistant Bringing the leading vector database to your cloud Getting started with llama-text-embed-v2 Natural Language Counterfactual Explanations for Graphs Using Large Language Models | Pinecone Easily build knowledgeable chat and agent-based applications in minutes with Pinecone Assistant, now generally available How to build an agentic, chat or RAG knowledge system using Pinecone Assistant Real-time RAG with Pinecone and Estuary Flow BigQuery to Pinecone in Real-Time with Estuary Flow Stravito Turns Market and Consumer Data Into Actionable Insights with Pinecone Inference | Pinecone Accelerate prototyping and development with Pinecone Local First-of-its-kind Pinecone Knowledge Platform to Power Best-in-class Retrieval for Customers Introducing integrated inference: Embed, rerank, and retrieve your data with a single API Strengthening security and increasing control with CMEK and API key roles Introducing Pinecone Rerank V0 Introducing cascading retrieval: Unifying dense and sparse with reranking From Idea to Action: How Pinecone Assistant Meaningfully Accelerates AI Business Building AI apps on Azure with Pinecone just got a lot easier Building a reliable, curated, and accurate RAG system with Cleanlab and Pinecone Four features of the Assistant API you aren't using - but should Deploying Pinecone with Infrastructure as Code (IaC) Streamlining CI/CD with Pinecone Local September 2024 Product Update Results of the Big ANN: NeurIPS'23 competition | Pinecone Introducing import from object storage for more efficient data transfer to Pinecone serverless Simplify, enhance, and evaluate RAG development with Pinecone Assistant, now in public preview Vectors and Graphs: Better Together August 2024 Product Update Pinecone Helps Deep Talk Deliver World-Class AI Assistants with Lower Engineering Overhead | Pinecone Assembled Delivers Better, Faster AI- Driven Support with Pinecone | Pinecone Llama 3.1 Agent using LangGraph and Ollama Build knowledgeable AI with Pinecone serverless, now generally available on Microsoft Azure Pinecone serverless is now generally available on Google Cloud, adding knowledge to AI assistants and other applications Accelerating Legal Discovery and Analysis with Pinecone and Voyage AI Bridging Dense and Sparse Maximum Inner Product Search | Pinecone Refine Retrieval Quality with Pinecone Rerank Introducing reranking to Pinecone Inference to simplify building accurate AI July 2024 Product Update Connect to Pinecone within your platform to enable a seamless AI development experience Introducing Pinecone API Versioning RAG Brag with Inkeep Co-Founder Nick Gomez LangGraph and Research Agents Introducing Pinecone Inference to streamline your AI workflow Build Privacy-aware AI software using Pinecone
Evolving Pinecone's architecture to meet the demands of Knowledgeable AI
Ram Sriharsha · 2025-02-25 · via Pinecone

Over the past year, we've seen a significant rise in demand for massive-scale knowledgeable AI applications as companies move from experimentation to production. In particular, this includes:

  • Recommender systems requiring 1000s of queries per second
  • Semantic search across over billions of documents
  • Agentic systems requiring millions of independent agents operating simultaneously

To tackle these scaled workloads head-on, we've continued to evolve our serverless architecture originally released in January 2024, distilling the insights gained from thousands of customers running on it in production. The result is a next-generation vector database that delivers significant advancements:

  • Predictable performance with the flexibility of a serverless architecture.
  • Highly reliable freshness and the capability to immediately reflect write operations for all workloads.
  • Cost effectiveness to run indexes with a large number of small namespaces.

Diverse workloads traditionally require different ANN algorithms

Traditionally, recommender systems have been treated as a “build once, serve many” form of indexing. Often, vector indexes for recommender workloads would be built in batch mode, taking hours. This means such indexes will be hours stale, but it also allows for heavy optimization of the serving index since it can be treated as static.

Graph algorithms like HNSW and Disk ANN are a good fit for this class of workloads. Graph-based indexes are challenging to update, but in this case, since index building happens offline, this isn’t such a big issue.

Furthermore, such indexes are often deployed on large machines with many cores (and high memory in the case of HNSW) and queries are batched as much as possible. Taken together, this configuration achieves the throughput for these relatively static workloads.

Semantic search workloads however are quite different. In these workloads, the corpora are generally far larger (100s of millions to billions of vectors) and require predictable low latency (O(100ms)) - but the throughput itself isn’t very high. They often employ heavy use of metadata filters. These workloads also care about freshness (i.e., whether Pinecone indexes reflect the most recent inserts and deletes) and see moderate to high update rates to their corpora.

This means that graph indexes aren’t best suited for these workloads. Nor are these workloads best served by vertically scaled machines.

Rather, they benefit from indexing techniques that can fall back to local SSD, as well as metadata filtering techniques that can leverage local SSD as much as possible.

And finally, agentic workloads have a very different characteristic compared to the above. They often have small to moderate sized corpora (often fewer than a million vectors) but a lot of namespaces (or tenants). Each tenant is small, and infrequently queried, but when a tenant is queried, the data is expected to be readily available and searched over.

Customers running workloads such as these expect certain features from Pinecone:

  • Highly-accurate vector search out of the box so they can focus on their business use cases. (They do not want to become vector search experts and start tuning knobs, choosing algorithms, etc.)
  • Freshness, elasticity, and the ability to ingest data. (They do not want to worry about hitting system limits, resharding, and resizing.)
  • Predictable, low latencies. (Or, at least, simple, intuitive ways to reason about latencies.)

So how can we handle all these diverse workloads well, given the above constraints?

Key Architectural Innovations

Log structured Indexing

Indexes in Pinecone serverless are composed of files. A collection of files can be thought of as an index (or more precisely a namespace, if you are using namespaces to partition indexes within Pinecone). These files are immutable, and are produced by an index builder.

The index builder has two competing objectives: in order to be fresh, it needs to quickly index data into a file and send it over to the read path for it to be served. At the same time, it needs to opportunistically look to spend more time indexing, so that it can produce a more optimal index for serving. In order to balance these priorities systematically, we employ a log structured indexing scheme that is similar to how log structured key value stores work.

Each file (or “slab” as we call it), can be thought of as logically consisting of the data, the index, the metadata index, and some book-keeping information.

The index type is self describing: this means that everything needed to decode the index and serve it is fully contained in the slab itself regardless of the type of the index.

Writes are first recorded in an in-memory structure called a memtable, which periodically flushes to blob storage, creating an L0 slab once it reaches a predefined size threshold. Since the goal is to make the L0 slabs indexed and queryable as soon as possible, we employ fast indexing techniques like scalar quantization or random projections at this level.

As the number of slabs reaches a size threshold, compaction kicks in. During compaction, smaller slabs are merged into larger slabs, and it is here that we build more computationally intensive partition- or graph-based indexes. By doing so, we amortize the cost of index building through the lifetime of serving those slabs. Compaction for large slabs is handled out of band and coordinated through the index builder.

Freshness

This approach automatically provides high freshness and paves the way to enable strongly consistent reads in the coming months since all reads are routed through the memtable on the index builder. The index builder is always aware of the latest manifest and upon reads, it issues a query to the executors to perform search over all the valid slabs and return the results, which are then merged to return the final candidates.

Users will also still be able to run in an eventually consistent mode where the queries are routed directly to the executors. This can be faster and also ensures high availability at the expense of a small amount of staleness where that tradeoff makes sense.

Predictable caching

The index portion of the slab is always cached between local SSD and memory. This allows us to serve queries immediately, without having to wait for a warmup period for cold queries.

The index builder sends a prewarming request to the executors upon generating slabs. The executors cache the slab, but reserve the right to evict the raw data if it is not touched for a while. Since the raw data is only needed to return the values and metadata, this means that even if the data isn’t cached it only incurs a small penalty (~250ms) until the data is cached again.

This typically happens after a few queries are run: during this time, the queries will show as cold in their response, and once the queries are operating on cached data, they will indicate as such in their response.

In the upcoming months, we will offer a mode called provisioned capacity, where you can reserve a certain amount of read units for a given index. In this mode, your data is guaranteed to be cached as long as it fits within those limits.

Provisioned capacity is useful when you have predictable workloads and you need eager caching and low tail latencies without incurring the overhead of cold queries.

Cost effective at high QPS

Given slabs have a tremendous amount of flexibility in what indexing can be applied to a given slab, this allows us to finally bring recommender workloads to serverless (where previously this was not possible) while enabling freshness for such workloads for the first time.

Given slabs are immutable, we can easily build any type of index including those that are graph-based, without having to worry about updates. Updates are handled through the log structured indexing scheme as mentioned above. Likewise, freshness is guaranteed by routing queries through the memtable. This allows us to optimize for high QPS workloads by building specific types of indexes that can be cached in memory and local SSD. Such slabs are deployed on specialized nodes with high CPU and memory bandwidth.

The immutability of slabs also makes it easy to replicate for throughput purposes. All of this happens automatically under the hood and needs no user intervention.

This allows us to bring the flexibility of serverless to recommendation workloads. You no longer have to worry about periodically building indexes, or hot swapping indexes in production, or what to do with the freshness lag. You get freshness, high QPS and high recall without any of the maintenance overheads.

Disk-based Metadata Filtering

Disk-based metadata filtering is another new feature in this update of our serverless architecture. Metadata filtering is a hard problem for vector databases. Both graph indexes and clustering-based indexes like IVF often struggle with metadata filters.

To power high-cardinality filtering use cases like access control lists, we've pioneered significant advancements in our single-stage filtering engine. We adapted the concept of bitmap indices commonly used in data warehouses and applied them to vector search. Every slab has a metadata folder, which can be thought of as a bitmap file per field. Low cardinality bitmaps are cached if they are frequently used (within a memory budget). High cardinality bitmaps are efficiently streamed from disk to intersect with the vector index itself in order to efficiently filter rows while scanning.

Some benchmarks

Our latest architecture provides better tail latencies while delivering better performance across the board for the same amount of cluster compute as the serverless system.

Disk-based metadata filters not only reduce memory usage and enable high cardinality filters to be applied efficiently, they also improve upon recall.

The Road Ahead

This iteration of our architecture represents a major step towards our commitment to deliver accurate, performant, and cost-effective retrieval for any scaled workload in production. But we're just getting started. On the immediate horizon, our next generation architecture enables us to:

  • Support seamless and cost-effective scaling to 1000+ QPS through provisioned read capacity
  • High performance sparse indexing for higher retrieval quality
  • Millions of namespaces per index to support massively multi-tenant use cases

We'll update you with more benchmarks and features during our upcoming Launch Week, March 17-21. Over the coming months, we'll be rolling out these architectural advancements to all serverless users. The best part? It's completely hands-off – simply keep building and let Pinecone handle the rest.