惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

The Cloudflare Blog
阮一峰的网络日志
阮一峰的网络日志
Martin Fowler
Martin Fowler
D
DataBreaches.Net
The GitHub Blog
The GitHub Blog
人人都是产品经理
人人都是产品经理
V
V2EX
爱范儿
爱范儿
PCI Perspectives
PCI Perspectives
T
Troy Hunt's Blog
Stack Overflow Blog
Stack Overflow Blog
www.infosecurity-magazine.com
www.infosecurity-magazine.com
SecWiki News
SecWiki News
cs.CV updates on arXiv.org
cs.CV updates on arXiv.org
The Hacker News
The Hacker News
小众软件
小众软件
雷峰网
雷峰网
D
Docker
NISL@THU
NISL@THU
cs.AI updates on arXiv.org
cs.AI updates on arXiv.org
腾讯CDC
B
Blog RSS Feed
C
CERT Recently Published Vulnerability Notes
cs.CL updates on arXiv.org
cs.CL updates on arXiv.org
U
Unit 42
有赞技术团队
有赞技术团队
P
Palo Alto Networks Blog
G
GRAHAM CLULEY
T
The Exploit Database - CXSecurity.com
T
Tailwind CSS Blog
S
Security @ Cisco Blogs
量子位
I
InfoQ
Application and Cybersecurity Blog
Application and Cybersecurity Blog
大猫的无限游戏
大猫的无限游戏
Schneier on Security
Schneier on Security
Help Net Security
Help Net Security
Latest news
Latest news
The Register - Security
The Register - Security
S
Securelist
W
WeLiveSecurity
A
Arctic Wolf
Security Latest
Security Latest
AWS News Blog
AWS News Blog
L
LINUX DO - 热门话题
S
Secure Thoughts
T
Tenable Blog
Know Your Adversary
Know Your Adversary
月光博客
月光博客
M
MIT News - Artificial intelligence

Pinecone

Pinecone Assistant: A Managed Knowledge Layer for Production AI Applications Multi-domain RAG in n8n: why one knowledge base is not enough Allspice Transforms the Culinary Experience with Semantic Search Powered by Pinecone | Pinecone Building RAG workflows in n8n: choosing the right Pinecone node Knowledge needs a meta-knowledge layer Garbage Day: How Pinecone Safely Deletes Billions of Objects at Scale When "Performance" Means Two Different Things Pinecone BYOC: Pinecone in your AWS, GCP, or Azure account, no vendor access True, Relevant, and Wrong: The Applicability Problem in RAG Use the Pinecone Plugin for Claude Code to develop AI Applications Faster Millions at Stake: How Melange's High-Recall Retrieval Prevents Litigation Collapse Powering High-stakes Patent Search at Scale: How Melange Built a Reliable AI System on Pinecone | Pinecone Pinecone Assistant Node in n8n: Turn Any Data Source Into Knowledge Pinecone Dedicated Read Nodes are now in Public Preview Inside Pinecone: Slab Architecture New Bulk Data Operations: Update, Delete, and Fetch by Metadata The Hidden Cost of Building: Lessons from Aquant Simplifying Vector Embeddings with Pinecone Integrated Inference Capabilities Pinecone joins Microsoft Marketplace as a Launch Partner GTM Engineering: Clay + Pinecone for AI-powered Sales Outbound Build an AI knowledge assistant with Google Docs and Pinecone Moving Pinecone forward with Ash Ashutosh as CEO and Edo spearheading our growing AI ambitions as Chief Scientist Pinecone Founder Edo Liberty to Spearhead Pinecone’s Growing AI Ambitions; Appoints Ash Ashutosh as CEO to Expand Vector Database Market Leadership Fast, Accurate Retrieval for Creators at Scale: Delphi’s Path Toward a Million Conversational Agents with Pinecone | Pinecone Announcing Pinecone Pioneers: A Program for Builders, Organizers, and Community Leaders What is Context Engineering? Chunking Strategies for LLM Applications Beyond the hype: Why RAG remains essential for modern AI Obviant Makes 30% More Accurate Defense Acquisition Recommendations Combining Sparse and Dense Retrieval with Pinecone | Pinecone Build more knowledgeable AI applications with new LLMs and greater control in Pinecone Assistant #NYTECHWEEK 2025 Retrieval-Augmented Generation (RAG) Accurate and Efficient Metadata Filtering in Pinecone’s Serverless Vector Database | Pinecone Terminal X AI Agents, Powered by Pinecone, Turn Complex Financial Data Into Production-grade Insights at Scale | Pinecone Aquant Delivers Scalable, Expert-level Service Intelligence with Pinecone | Pinecone Cascading retrieval with multi-vector representations: balancing efficiency and effectiveness Vector databases aren't just for large-scale enterprise AI Unveiling DIME: Reproducibility, Scalability, and Formal Analysis of Dimension Importance Estimation for Dense Retrieval | Pinecone Fast and Effective Early Termination for Simple Ranking Functions | Pinecone Domain-specific AI Agents at Scale: CustomGPT.ai Serves 10,000+ Customers with Pinecone | Pinecone Using Pinecone asynchronously with FastAPI A Flexible Resource for Top-Weighted Comparisons Between Sets and Rankings | Pinecone Build secure, scalable agentic AI workflows with Rubrik Annapurna and Pinecone Tool up: Pinecone’s first MCP servers are here Add context to your agent with Pinecone Assistant MCP remote server E2Rank: Efficient and Effective Layer-wise Reranking | Pinecone ColBERT-serve: Efficient Multi-Stage Memory-Mapped Scoring | Pinecone Efficient Constant-Space Multi-Vector Retrieval | Pinecone How Vanguard Worked with Pinecone to Boost Customer Support with Faster Calls and 12% More Accurate Responses | Pinecone Pinecone Named to Fast Company's Annual List of the World's Most Innovative Companies of 2025 Launch Week: Pinecone for agents, search, recommendations, and more Optimizing Pinecone for agents (and more) Retrieval Inference for scale and performance How 1up Turns Sales Reps Into Product Experts with Pinecone | Pinecone Don’t be dense: Launching sparse indexes in Pinecone Unlock High-Precision Keyword Search with pinecone-sparse-english-v0 Evolving Pinecone's architecture to meet the demands of Knowledgeable AI Pinpoint references faster with citation highlights in Pinecone Assistant Bringing the leading vector database to your cloud Getting started with llama-text-embed-v2 Natural Language Counterfactual Explanations for Graphs Using Large Language Models | Pinecone Easily build knowledgeable chat and agent-based applications in minutes with Pinecone Assistant, now generally available How to build an agentic, chat or RAG knowledge system using Pinecone Assistant Real-time RAG with Pinecone and Estuary Flow BigQuery to Pinecone in Real-Time with Estuary Flow Stravito Turns Market and Consumer Data Into Actionable Insights with Pinecone Inference | Pinecone Accelerate prototyping and development with Pinecone Local First-of-its-kind Pinecone Knowledge Platform to Power Best-in-class Retrieval for Customers Introducing integrated inference: Embed, rerank, and retrieve your data with a single API Strengthening security and increasing control with CMEK and API key roles Introducing Pinecone Rerank V0 Introducing cascading retrieval: Unifying dense and sparse with reranking From Idea to Action: How Pinecone Assistant Meaningfully Accelerates AI Business Building AI apps on Azure with Pinecone just got a lot easier Building a reliable, curated, and accurate RAG system with Cleanlab and Pinecone Four features of the Assistant API you aren't using - but should Deploying Pinecone with Infrastructure as Code (IaC) Streamlining CI/CD with Pinecone Local September 2024 Product Update Results of the Big ANN: NeurIPS'23 competition | Pinecone Introducing import from object storage for more efficient data transfer to Pinecone serverless Simplify, enhance, and evaluate RAG development with Pinecone Assistant, now in public preview Vectors and Graphs: Better Together August 2024 Product Update Pinecone Helps Deep Talk Deliver World-Class AI Assistants with Lower Engineering Overhead | Pinecone Assembled Delivers Better, Faster AI- Driven Support with Pinecone | Pinecone Llama 3.1 Agent using LangGraph and Ollama Build knowledgeable AI with Pinecone serverless, now generally available on Microsoft Azure Pinecone serverless is now generally available on Google Cloud, adding knowledge to AI assistants and other applications Accelerating Legal Discovery and Analysis with Pinecone and Voyage AI Bridging Dense and Sparse Maximum Inner Product Search | Pinecone Refine Retrieval Quality with Pinecone Rerank Introducing reranking to Pinecone Inference to simplify building accurate AI July 2024 Product Update Connect to Pinecone within your platform to enable a seamless AI development experience Introducing Pinecone API Versioning RAG Brag with Inkeep Co-Founder Nick Gomez LangGraph and Research Agents Introducing Pinecone Inference to streamline your AI workflow Build Privacy-aware AI software using Pinecone
RAG with Access Control
Sohan Maheshwar · 2026-01-08 · via Pinecone

This Learn article was written by Sohan Maheshwar, Lead Developer Advocate at AuthZed.

Securing Knowledge Retrieval using SpiceDB and ReBAC

Retrieval-Augmented Generation (RAG) pipelines are powerful tools for grounding large language models (LLMs) in private or domain-specific knowledge. By fetching relevant documents from a vector database and including them in the model’s prompt, RAG enables more accurate and context-aware answers.

However, this capability introduces a serious risk: information leakage. If different users have different levels of access to data, as they do in most real-world systems, your RAG pipeline must enforce those access boundaries. In fact, OWASP lists Sensitive Information Disclosure, Excessive Agency & Vector and Embedding Weaknesses in their list of Top 10 Risks for Large Language Model Applications.

That’s where access control comes in. Instead of trusting the RAG system to “just do the right thing,” we can integrate an authorization layer that determines which resources a given identity can access. In this post, we’ll explore how to achieve that using SpiceDB, an open-source, Zanzibar-inspired permission system.

We’ll cover how SpiceDB works, how to model permissions, and how to apply access control both before and after retrieval in a RAG pipeline built with Pinecone and OpenAI embeddings.

Before we get to the implementation, let’s review two foundational concepts: authentication and authorization.

  • Authentication: Verifying who a user is. Usually handled through credentials, OAuth, or an identity provider.
  • Authorization: Deciding what that user can do once authenticated — which resources they can read, write, or modify.

Several access control paradigms exist:

  • ACL (Access Control Lists): Simple user-resource mappings.
  • RBAC (Role-Based Access Control): Permissions granted through roles like “admin” or “editor.”
  • ABAC (Attribute-Based Access Control): Decisions based on attributes such as department, geography, or clearance.
  • ReBAC (Relationship-Based Access Control): Permissions determined by relationships between users and resources.

For large, dynamic, and context-rich applications (such as RAG pipelines) ReBAC provides the flexibility and scalability we need. It models access as a graph of relationships rather than hard-coded rules.

Enter Google Zanzibar

Google Zanzibar is the internal authorization system that Google built to manage permissions across all their products and services. Think of it as the system that decides whether you can view a shared Google Doc, edit a file in Google Drive, or access a specific Google Cloud resource. Rather than each Google product implementing its own permission system, they all use Zanzibar as a shared service.

Zanzibar systems shine when the requirement is for:

  • Low-latency
  • High-throughput authorization checks
  • Global consistency of relationship data
  • Composable and hierarchical permission models

Turns out, RAG systems typically have these requirements. In addition, embeddings need to retain the permissions from where the data was originally sourced. This means that the permission system needs to be powerful and flexible enough to model all of the source systems' permission systems. These requirements make ReBAC & a Zanzibar-like system the perfect way to implement fine-grained permissions for RAG.

SpiceDB

While Zanzibar is an internal system at Google, there are open source implementations based on the whitepaper that Google released in 2019. Among those, SpiceDB is the most scalable and consistent open-source implementation of Google’s Zanzibar authorization model, used by businesses such as OpenAI, Workday, Turo, Netflix and more.

SpiceDB stores access relationships as a graph, where nodes represent entities (users, groups, documents) and edges represent relationships (like “viewer,” “editor,” or “owner”). Fundamentally, authorization logic can be reduced to asking a single question:

Is this actor allowed to perform this action on this resource?

In SpiceDB parlance, this actor and this resource are both Objects and this action is a Permission or Relation. Here’s a Google Docs style example where a user can be either a reader or a writer of a document. A reader can only read the document, whereas a writer can read and write the document.

You can represent this use case using a schema like this:

definition user {}

definition document {
    relation reader: user
    relation writer: user

    permission can_read = reader + writer
    permission can_write = writer
}

This defines:

  • user and document object types
  • Relations (reader, writer)
  • Derived permissions (can_read, can_write)

When a user requests access to a document, SpiceDB answers questions like:

“Can user:alice perform can_read on document:doc1?”

It evaluates the relationship graph in microseconds, allowing real-time authorization checks at massive scale.

In our RAG pipeline, we’ll use this to ensure that each document fetched from the vector store is checked against the access graph before it’s used in the model’s context.

Access Control for RAG Applications - Simulating a RAG Pipeline

We’ll now simulate a simple RAG setup using:

  • OpenAI and Langchain for embeddings,
  • Pinecone for vector similarity search, and
  • SpiceDB for authorization.

This article distills the concepts of RAG and Authorization to a simple use case: we have a user ‘Kim’ that has access to two documents with IDs 123 and 456.

We query an LLM for information from one of the documents. We then remove Kim's permissions to view one of the documents and make the same query. If all goes well, the information should not be available to them.

1. Create a Pinecone Index

Let’s define a Pinecone serverless index that we’ll use to store our vectors:

pc = Pinecone(api_key=PINECONE_API_KEY)

index_name = "documents"
namespace_name = "authzed"

pc.create_index(
    name=index_name,
    dimension=1024,
    metric="cosine",
    spec=ServerlessSpec(
        cloud="aws",
        region="us-east-1"
    )
)

We are simulating a real-world RAG (retrieval-augmented generation) scenario by embedding 2 documents.

In this example, we also specify metadata like article_id to track which article the chunk comes from. The article_id is important for linking embeddings to objects that users are authorized on.

# Create a Document object that specifies our made up documents and specifies the document_id as metadata.

documents = [
    Document(
        page_content="The number of paid holidays is 50 per year",
        metadata={"article_id": "123"}
    ),
    Document(
        page_content="The revenue for Q4 2025 is one billion dollars!",
        metadata={"article_id": "456"}
    )
]


# Initialize LangChain embeddings
embeddings = PineconeEmbeddings(
    model="multilingual-e5-large",
    pinecone_api_key=PINECONE_API_KEY
)

# Create vector store and upsert both documents
docsearch = PineconeVectorStore.from_documents(
    documents=documents,
    index_name=index_name,
    embedding=embeddings,
    namespace=namespace_name
)

2. Installing SpiceDB

There are multiple ways of running an instance of SpiceDB after installing it.

To run it locally you can run the binary:

spicedb serve --grpc-preshared-key "t_your_token_here_1234567deadbeef"

Or in a local Docker container:

docker run --rm -p 50051:50051 authzed/spicedb serve --grpc-preshared-key "t_your_token_here_1234567deadbeef"

You could also use the AuthZed Cloud to create a SpiceDB Permissions System in the cloud for production purposes.

This article will assume a local instance of SpiceDB.

For our RAG pipeline, the schema below defines two object types, user and article. Users can relate to an article as a viewer and any user who is related to the article as a viewer has view permissions on the document.

def make_spicedb_client() -> Client:
    # For TLS environments, replace with bearer_token_credentials(...).
    return Client(
        target=SPICEDB_ENDPOINT,
        credentials=insecure_bearer_token_credentials(SPICEDB_TOKEN),
    )

client = make_spicedb_client()
print('SpiceDB client ready:', isinstance(client, Client))

SCHEMA = """definition user {}

definition article {
    relation viewer: user

    permission view = viewer
}"""

try:
    resp = await client.WriteSchema(WriteSchemaRequest(schema=SCHEMA))
except Exception as e:
    print(f"Write schema error: {type(e).__name__}: {e}")

3. Write Relationships to SpiceDB

Now, we write relationships to SpiceDB that specify that Kim is a viewer of document 123 and 456.

After these relationships are written, any permission checks to SpiceDB will reflect that Kim can view documents 123 and 456. In production, when a user creates a new document, a relationship is created between the user and that document.

try:
    resp = await (client.WriteRelationships(
        WriteRelationshipsRequest(
            updates=[
                RelationshipUpdate(
                    operation=RelationshipUpdate.Operation.OPERATION_TOUCH,
                    relationship=Relationship(
                        resource=ObjectReference(object_type="article", object_id="123"),
                        relation="viewer",
                        subject=SubjectReference(
                            object=ObjectReference(
                                object_type="user",
                                object_id="kim",
                            )
                        ),
                    ),
                ),
                RelationshipUpdate(
                    operation=RelationshipUpdate.Operation.OPERATION_TOUCH,
                    relationship=Relationship(
                        resource=ObjectReference(object_type="article", object_id="456"),
                        relation="viewer",
                        subject=SubjectReference(
                            object=ObjectReference(
                                object_type="user",
                                object_id="kim",
                            )
                        ),
                    ),
                ),
            ]
        )
    ))
except Exception as e:
    print(f"Write relationships error: {type(e).__name__}: {e}")

Authorization Techniques using ReBAC

There are two techniques of fine-grained authorization in RAG using SpiceDB:

  1. Pre-filter
  2. Post-Filter.

Let’s take a look at how both methods would work and when you would choose to use one over the other.

1. Pre-Filter Authorization

In this method we'll first query SpiceDB for a list of documents that Kim is allowed to view using the LookupResources API. Only the list of documents that the user can access is embedded and sent to the vector database ensuring that the user’s query to the LLM includes only documents they have permission to view.

Here is a high-level architecture diagram of the Pre-Filter method:

High-level architecture diagram of the Pre-Filter method using SpiceDB

subject = SubjectReference(
    object=ObjectReference(
        object_type="user",
        object_id="kim",
    )
)

def lookupArticles():
    return client.LookupResources(
        LookupResourcesRequest(
            subject=subject,
            permission="view",
            resource_object_type="article",
        )
    )
try:
    resp = lookupArticles()

    authorized_articles = []

    async for response in resp:
            authorized_articles.append(response.resource_object_id)
except Exception as e:
    print(f"Lookup error: {type(e).__name__}: {e}")

print("Article IDs that Kim is authorized to view:")
print(authorized_articles)

The result of the lookupResources call looks like this:

Article IDs that Kim is authorized to view:
['123', '456']

We can now issue a prompt to GPT-5, enhanced with relevant data that the user is authorized to access. This ensures that the response is based on information the user is permitted to view.

We initialize a LangChain object to create an OpenAI chat model (this could be swapped out for any LLM). We also initialize a LangChain object for a Pinecone index with the OpenAI embeddings model. The retriever is a filter that restricts the search to authorized documents. We also add some basic prompt engineering to ensure the answer from the LLM is restricted to the context passed.

# Define the ask function
def ask():
    # Initialize a LangChain object for an OpenAI chat model.
    llm = ChatOpenAI(
        openai_api_key=OPENAI_API_KEY,
        model="gpt-5-nano-2025-08-07",
        temperature=1
    )

    # Initialize a LangChain object for a Pinecone index with an OpenAI embeddings model.
    knowledge = PineconeVectorStore.from_existing_index(
        index_name=index_name,
        namespace=namespace_name,
        embedding=OpenAIEmbeddings(
            openai_api_key=OPENAI_API_KEY,
            dimensions=1024,
            model="text-embedding-3-large"
        )
    )

    # Initialize a retriever
    retriever=knowledge.as_retriever(
            search_kwargs={
            "filter": {
                "article_id":
                    {"$in": authorized_articles},
            },
        }
    )

    # Initialize a string prompt template that let's us add context and a question.
    prompt = ChatPromptTemplate.from_template("""Answer the question below using the context:

    Context: {context}

    Question: {question}

    Answer: """)

    retrieval =  RunnableParallel(
        {"context": retriever, "question": RunnablePassthrough()}
    )

    chain = retrieval | prompt | llm | StrOutputParser()

    question = """How many paid holidays do I have this year"""

    print("Prompt: \n")
    print(question)
    print(chain.invoke(question))

#invoke the ask function
ask()

The LLM returns something like this:

Prompt: 

How many paid holidays do I have this year
50 paid holidays.

You can also generate a summary of all the articles that Kim is authorized to view.

async def summarize_accessible_articles(user_id: str):

    # 1️⃣ Lookup articles
    subject = SubjectReference(
        object=ObjectReference(object_type="user", object_id=user_id)
    )
    response = client.LookupResources(
        LookupResourcesRequest(
            subject=subject,
            permission="view",
            resource_object_type="article",
        )
    )
    authorized_articles = [res.resource_object_id async for res in response]
    print(f"🔍 {user_id} can view articles: {authorized_articles}")

    if not authorized_articles:
        return "❌ No accessible articles."

    # 2️⃣ Setup LangChain retriever w/ filter
    knowledge = PineconeVectorStore.from_existing_index(
        index_name=index_name,
        namespace=namespace_name,
        embedding=OpenAIEmbeddings(
            openai_api_key=OPENAI_API_KEY,
            dimensions=1024,
            model="text-embedding-3-large"
        )
    )

    retriever = knowledge.as_retriever(
        search_kwargs={
            "filter": {"article_id": {"$in": authorized_articles}},
            "k": 100  # Ensure we get all matches
        }
    )

    docs = await retriever.ainvoke("Give me all the contents to summarize")

    if not docs:
        return "❌ No content found."

    combined_text = "\n\n".join([d.page_content for d in docs])

    # 3️⃣ Summarize using OpenAI
    summary_prompt = (
        "You are an AI assistant. Based ONLY on the following articles, "
        "generate a concise summary of their contents. Do not use any outside knowledge.\n\n"
        + combined_text
        + "\n\nSummary:"
    )

    openai_client = AsyncOpenAI(api_key=OPENAI_API_KEY)
    chat_response = await openai_client.chat.completions.create(
        messages=[{"role": "user", "content": summary_prompt}],
        model="gpt-5-nano-2025-08-07",
        temperature=1
    )

    return chat_response.choices[0].message.content

# Make the call to summarize_accessible_articles
summary = await summarize_accessible_articles("kim")
print("📄 Summary of accessible articles:")
print(summary)

And the resulting output:

🔍 Kim can view articles: ['123', '456']
📄 Summary of accessible articles:
Summary: There are 50 paid holidays per year, and Q4 2025 revenue is one billion dollars.

A Change in Permission

Let’s make a request when the user is NOT authorized to view the necessary contextual data. I.e let's see what happens when Kim is not authorized to view the document.

We can simulate this by deleting the relationship between Kim as a viewer to document 456. In production this would be an API call to the RelationshipUpdate API.

try: 
    resp = await client.WriteRelationships(
        WriteRelationshipsRequest(
            updates=[
                RelationshipUpdate(
                    operation=RelationshipUpdate.Operation.OPERATION_DELETE,
                    relationship=Relationship(
                        resource=ObjectReference(object_type="article", object_id="123"),
                        relation="viewer",
                        subject=SubjectReference(
                            object=ObjectReference(
                                object_type="user",
                                object_id="kim",
                            )
                        ),
                    ),
                ),
            ]
        )
    )
except Exception as e:
    print(f"Write relationships error: {type(e).__name__}: {e}")

To confirm this update in permissions, run the LookupResources method again:

try:
        resp = lookupArticles()

        authorized_articles = []

        async for response in resp:
                authorized_articles.append(response.resource_object_id)
except Exception as e:
    print(f"Lookup error: {type(e).__name__}: {e}")

print("Documents that Kim can view:")
print(authorized_articles)

And the output:

Documents that Kim can view:
['456']

Now, we can run our earlier query again:

With an output:

Prompt: 

How many paid holidays do I have this year
The provided context does not specify any paid holidays; it only mentions Q4 2025 revenue (one billion dollars). If you need to know your paid holidays, please share your company’s holiday policy or HR information, or check your employee handbook/HR portal.

We no longer receive a completion that answers our question because Kim is no longer authorized to view the document that contains the context about the number of paid holidays in the year.

2. Post-Filter Method

We just completed the pre-filter method where we queried SpiceDB for all the documents that Kim was authorized to view. An alternate approach is to use the Post-Filter method where a CheckPermissionRequest is performed on every document ID that the vector database returns. The list of authorized documents is then passed on to the LLM for a response to the query.

Here is a high-level architecture diagram of the post-filter method:

High-level architecture diagram of the post-filter method using SpiceDB

Let's restore Kim's permissions to view document 123:

try: 
    resp = await client.WriteRelationships(
        WriteRelationshipsRequest(
            updates=[
                RelationshipUpdate(
                    operation=RelationshipUpdate.Operation.OPERATION_TOUCH,
                    relationship=Relationship(
                        resource=ObjectReference(object_type="article", object_id="123"),
                        relation="viewer",
                        subject=SubjectReference(
                            object=ObjectReference(
                                object_type="user",
                                object_id="kim",
                            )
                        ),
                    ),
                ),
            ]
        )
    )
except Exception as e:
    print(f"Write relationships error: {type(e).__name__}: {e}")

Checking For Permissions

Define the method that gets the article_id for all documents and checks whether the user has permissions for each article. Compare and contrast this with the Pre-Filter method where we performed a lookup to get a list of documents that the user had access to.

async def filter_docs_with_spicedb(docs: List):
    filtered_docs = []
    for doc in docs:
        article_id = doc.metadata.get("article_id")
        resp = await client.CheckPermission(
            CheckPermissionRequest(
                subject=SubjectReference(
                    object=ObjectReference(
                        object_type="user",
                        object_id="kim",
                    ),
                ),
                resource=ObjectReference(
                    object_type="article",
                    object_id=str(article_id),
                ),
                permission="view",
            )
        )
        if resp.permissionship == CheckPermissionResponse.PERMISSIONSHIP_HAS_PERMISSION:
            filtered_docs.append(doc)
        
    return filtered_docs

All that's left is to build a Langchain object. This snippet sets up a retriever to fetch relevant documents, then applies a post-filter using SpiceDB to ensure only documents the user is authorized to view are included.

The RunnableLambda allows you to wrap a custom Python function (such as your authorization filter) so it can be used as a step in the LangChain pipeline. RunnablePassthrough simply passes its input through unchanged, making it useful for forwarding data (like the user's question) to the next step in the chain.

# Build the LangChain object
retriever = docsearch.as_retriever(search_kwargs={"k": 4})
llm = ChatOpenAI(api_key=OPENAI_API_KEY, 
                 model="gpt-5-nano-2025-08-07", 
                 temperature=1)

prompt = ChatPromptTemplate.from_messages([
    ("system", "You answer strictly from the provided context. If insufficient, say so."),
    ("human", "Question: {question}\n\nContext:\n{context}")
])

# Combine: retrieve → post-filter → prompt → LLM
graph = (
    RunnableParallel(
        {
            "context": retriever | RunnableLambda(filter_docs_with_spicedb),
            "question": RunnablePassthrough(),
        }
    )
    | prompt
    | llm
    | StrOutputParser()
)

print("✅ Retrieval + chain wired up")

Run this code to ask the LLM about some data in document 123. Since Kim does have permission to view this document, you should see the correct response.

question = "How many holidays can I take this year?"
result = await graph.ainvoke(question) 
print(result)

The LLM responds with:

You can take up to 50 paid holidays this year.

Let's repeat the same steps from the pre-filter technique by removing Kim's permission to view the document and then asking the same question. Since Kim doesn't have permission to view this document the LLM isn't able to provide an answer.

try: 
    resp = await client.WriteRelationships(
        WriteRelationshipsRequest(
            updates=[
                RelationshipUpdate(
                    operation=RelationshipUpdate.Operation.OPERATION_DELETE,
                    relationship=Relationship(
                        resource=ObjectReference(object_type="article", object_id="123"),
                        relation="viewer",
                        subject=SubjectReference(
                            object=ObjectReference(
                                object_type="user",
                                object_id="kim",
                            )
                        ),
                    ),
                ),
            ]
        )
    )
except Exception as e:
    print(f"Write relationships error: {type(e).__name__}: {e}")


question = "How many holidays can I take this year?"
result = await graph.ainvoke(question) 
print(result)

The response:

Insufficient information. The provided document only mentions Q4 2025 revenue and has no details about holidays or leave policies. If you share your company’s holiday/leave policy or HR guidelines, I can help interpret it.

How do you choose between using the pre-filter or the post-filter technique? It Depends™️.

Typically, if you have a high positive hit-rate of documents from your vector database, a post-filter approach works well. Conversely, if you have a large corpus of documents in your RAG pipeline and a low positive hit-rate, the pre-filter approach is more efficient. It’s important to note that the lookupResources call is computationally more expensive than a checkPermission call. Also, you can perform multiple permission checks in a single call with CheckBulkPermissionsRequest. The general advice is to look at your SpiceDB metrics to determine which approach works best.

Takeaways

Access control is fundamental to building trustworthy AI systems. In RAG pipelines, where content retrieval directly influences model output, fine-grained authorization must be enforced at every layer.

By integrating SpiceDB into your RAG pipeline, you get:

  • Centralized authorization modeled as relationships.
  • Low-latency ReBAC checks optimized for high throughput.
  • Clear separation of concerns between search, embedding, and access control.

Congratulations! You learned how to secure your RAG pipelines with fine-grained authorization using SpiceDB.

OpenAI uses SpiceDB and AuthZed Dedicated to secure 37 Billion documents for 5 Million users who use ChatGPT Connectors. Read more about it here: https://authzed.com/customers/openai.