惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

博客园 - 叶小钗
O
OpenAI News
V
V2EX
大猫的无限游戏
大猫的无限游戏
博客园 - 聂微东
S
Schneier on Security
C
CXSECURITY Database RSS Feed - CXSecurity.com
小众软件
小众软件
L
LINUX DO - 热门话题
C
Cybersecurity and Infrastructure Security Agency CISA
博客园 - Franky
Security Latest
Security Latest
S
SegmentFault 最新的问题
Project Zero
Project Zero
Spread Privacy
Spread Privacy
K
Kaspersky official blog
J
Java Code Geeks
V
Vulnerabilities – Threatpost
C
Cisco Blogs
C
CERT Recently Published Vulnerability Notes
月光博客
月光博客
T
The Exploit Database - CXSecurity.com
L
Lohrmann on Cybersecurity
人人都是产品经理
人人都是产品经理
博客园 - 三生石上(FineUI控件)
Scott Helme
Scott Helme
WordPress大学
WordPress大学
量子位
T
Threat Research - Cisco Blogs
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻
宝玉的分享
宝玉的分享
Hugging Face - Blog
Hugging Face - Blog
AWS News Blog
AWS News Blog
Help Net Security
Help Net Security
Application and Cybersecurity Blog
Application and Cybersecurity Blog
Simon Willison's Weblog
Simon Willison's Weblog
S
Secure Thoughts
博客园 - 【当耐特】
cs.CV updates on arXiv.org
cs.CV updates on arXiv.org
V
Visual Studio Blog
Last Week in AI
Last Week in AI
T
Tailwind CSS Blog
腾讯CDC
Cyberwarzone
Cyberwarzone
IT之家
IT之家
GbyAI
GbyAI
Exploit-DB.com RSS Feed
Exploit-DB.com RSS Feed
云风的 BLOG
云风的 BLOG
T
Troy Hunt's Blog
D
Docker

Analytics Vidhya

Handling Imbalanced Classification: What Works Better Than SMOTE GPT-5.6 Is Here: Sol, Terra, and Luna Loop Engineering for AI Agents: How /loop is Changing AI Workflows DeepSeek DSpark: The Speculative Decoding Trick Behind 400% Faster LLM Modern VLMs Explained: How GPT-4o, Gemini, Claude Vision, and Qwen-VL Work YOLO26 Tutorial: Object Detection, Pose Estimation & More Large Action Models (LAMs) vs Agentic LLMs: What's the Real Difference? Claude Sonnet 5: The Fable 5 at Home The Best $20 AI Plan: ChatGPT Plus vs Claude Pro vs Gemini Pro GraphRAG vs Vector RAG: Which Retrieval Method is Best? Using AI When You Don’t Trust AI The Self-Improving Loop in AI Agents: Architecture, Benefits, and How it Outperforms Traditional Agent Workflows Harness-1: The 20B Retrieval Subagent That Beats GPT-5.4 at Search Sakana Fugu: Multi-Agent System as a Model Claude's Hidden Art Skill: Making Illustrations With Code System Design for ML Interviews: 10 Real Problems Walked Through Most People Use ChatGPT Wrong: 10 Features and Tips That Changed How I Work OpenAI Just Launched 3 Free AI Courses with Certificates Autoregressive Models: Predicting the Future Using the Past Gemini Omni: AI Video Generation Inside Gemini DiffusionGemma: Google’s Diffusion-Based Open Model for Faster Text Generation Top 10 AI Engineering Tools Everyone is Using in 2026 I Tested Claude Fable 5: Can Anthropic’s Newest AI Deliver on the Hype? Prophet vs NeuralProphet vs TimeGPT vs Chronos: A Practical Comparison Build an Emergency Helpline Voice Agent with LangChain Choosing the Right Vector Database for RAG and AI Applications Google Gemma 4 12B: Architecture, Benchmarks, Access, and Hands-on Guide for Developers How to Choose the Right AI Model for Your Needs Agent Observability with LangSmith, Langfuse, and Arize: A Hands-On Comparison How to Use Claude Managed Agents? Google AI Studio vs Gemini App: What’s the Difference? AI Workflows for Sales Teams: Prospect Research, Lead Qualification, and CRM Updates on Autopilot Using LangGraph 25 Most Influential AI Pioneers to Meet at DataHack Summit 2026 Claude Opus 4.8: A Smarter Model in the Right Direction PySpark Optimization: 12 Proven Techniques to Speed Up Your Spark Jobs 10 Everyday Tasks You Can Automate with AI Today (With n8n Templates) Google Antigravity 2.0: The Full Developer Guide (I/O 2026) Build a Claude Cowork-Like Browser Agent Using Playwright MCP and Claude Desktop Pandas vs Polars vs DuckDB: Which Library Should You Choose? Qwen3.7-Max: Alibaba’s New Agent-First LLM for Coding, Reasoning, and Long-Horizon AI Workflows The Biggest Announcements from Google I/O 2026 Top 9 AI Events and Conferences in 2026 that you Must Attend Gemini 3.5 Flash: Frontier Intelligence with Speed Kimi WebBridge: Hands-on Guide to Kimi’s Browser Extension for AI Agents 40 Advanced SQL Window Functions Every Data Scientist Must Know(with examples) Top 10 AI Research Papers of 2025 6 Steps to Crack GenAI Case Study Interviews (With Real Examples) OpenAI Omni Moderation: How to Filter Text & Images for Free DataHack Summit 2026: You Just Cannot Skip This AI Event of the Year OpenAI’s New API Voice Models Will Change the Way You Use AI Hermes Agent Guide: What is it and How to Use it? Top 10 LLM Research Papers of 2026 Agent Memory Patterns in Cognitive Science and AI Systems 10 AI Agents Every AI Engineer Must Build (with GitHub Samples) 23 Tips for Smart Claude Code Token Saving and Workflow Optimization Feature Engineering with LLMs: Techniques & Python Examples ChatGPT is Now Inside Excel and Google Sheets: Here is How to Use it Gemini API File Search: The Easy Way to Build RAG Top 10 Open-Source Libraries to Fine-Tune LLMs Locally ML Intern in Practice: From Prompt to a Shipped Hugging Face Model 15+ Solved Agentic AI Projects with Github Links How People are Figuring Out Life With Claude MemPalace Explained: Building Long-Term Memory for AI Agents Beyond RAG Grok Voice Think Fast 1.0: Build Voice AI Agents That Actually Think Compressing LSTM Models for Retail Edge Deployment: A Practical Comparison MCP vs Agent Skills: Different Altogether GPT 5.5 vs Opus 4.7: Which is the Best AI Model Today? What is Agentic AI? Claude Code vs Codex: A Detailed Terminal Agent Comparison Google Deep Research Max: Build Autonomous AI Research Agents in Minutes Meta Muse Spark Review: Is It Worth the Hype? ChatGPT Images 2.0 vs Nano Banana 2: Which is Better? Cursor V3 Explained: The AI Coding Agent That’s Replacing Traditional IDEs in 2026 DeepSeek-V4: The Most Powerful Open-Source Model Ever Is GPT Image 2 the Best Image Generation Model? Token Economics: Why AI is Getting “Cheaper” From Idea to Output: Claude Does the Design Work Opus 4.7 vs Opus 4.6: Should You Switch? Build Human-Like AI Voice App with Gemini 3.1 Flash TTS How to Structure a Claude Code Project that Thinks Like an Engineer Gemma 4 Tool Calling Explained: Build AI Agents with Function Calling (Step-by-Step Guide) Anthropic Launches Claude Opus 4.7 For “Most Difficult Tasks” Top 28 Claude Shortcuts that will 10X your Speed GPT-5.4-Cyber: Why OpenAI is Keeping its Most Powerful Model Under Lock and Key Google AI Studio Guide: Every Feature Explained Mastering Deep Agents: Context Engineering that Actually Works 21 Computer Vision Projects from Beginner to Advanced (2026 Guide) Excel 101: Excel Agent Mode Explained MiniMax M2.7 Goes Open-Weight to Let You Run Agents Locally Top 10 Gemma 4 Projects That Will Blow Your Mind GLM-5.1: Architecture, Benchmarks, Capabilities & How to Use It Understanding BERTopic: From Raw Text to Interpretable Topics From Karpathy’s LLM Wiki to Graphify: AI Memory Layers are Here 10 Most Important AI Concepts Explained Simply Project Glasswing is World’s Most Powerful AI in Action How to Run Gemma 4 on Your Phone Without Internet: A Hands-On Guide Running Claude Code for Free with Gemma 4 and Ollama LLM Wiki Revolution: How Andrej Karpathy’s Idea is Changing AI Rethinking Enterprise Search: How Cortex Search Turns Data into Business Impact Google’s Gemma 4: Is it the Best Open-Source Model of 2026?
OKF: Redefining Knowledge Bases for AI Agents
Shaik Hamzah · 2026-07-07 · via Analytics Vidhya

In June 2026, Google introduced the Open Knowledge Format (OKF), an open specification for how AI agents organise and exchange knowledge. An OKF bundle is just Markdown files, lightweight YAML metadata, and links between concepts, yet it challenges the assumption that every AI application needs embeddings and vector databases.

Because the knowledge base is plain text, it can be version-controlled in Git and navigated by following links rather than retrieving disconnected chunks. In this article, we’ll explore how OKF works and when it beats a traditional retrieval pipeline.

Table of contents

  • Why Traditional RAG Has Limitations
  • What is the Open Knowledge Format (OKF)?
  • Structure of an OKF Bundle
  • A Typical OKF Folder Structure
  • Building an OKF Bundle
  • Creating an OKF Concept File
  • Understanding the YAML Metadata
  • Linking Related Concepts
  • Another Example: Hospital Metric
  • How AI Agents Traverse an OKF Bundle
  • Why This Works Well for AI Agents
  • Where RAG Still Excels
  • Hybrid Knowledge Architecture: Combining OKF and RAG
  • Conclusion
  • Frequently Asked Questions

Why Traditional RAG Has Limitations

Over the past few years, Retrieval-Augmented Generation (RAG) has become the standard approach for providing external knowledge to Large Language Models. Instead of relying solely on the model’s training data, RAG retrieves relevant information from external documents during inference. A typical pipeline looks something like this:

Traditional RAG Pipeline

This approach works remarkably well for searching millions of documents. By comparing the semantic meaning of embeddings instead of exact keywords, RAG allows an AI system to answer questions using information that was never part of the model’s original training data.

However, there is an important trade-off. Before a document can be indexed, it must first be divided into smaller chunks. While chunking improves retrieval efficiency, it also breaks apart the original structure of the document. Relationships that were naturally connected inside a single document become distributed across multiple independent chunks.

Consider the following hospital protocol.

# Patient Admission Policy

Patients arriving through the Emergency Department must complete an initial triage before admission.

## Admission Requirements

- Valid patient identification
- Initial clinical assessment completed
- Emergency cases receive immediate priority

## Recording and Bed Allocation

Patient information is recorded in the Electronic Health Record (EHR) system before a bed is assigned.

Bed allocation follows the Bed Occupancy guidelines maintained by the Operations team.

A typical RAG pipeline may split this document into several smaller chunks before indexing.

1st Chunk

Patients arriving through the Emergency Department must complete an initial triage.

Admission Requirements:

- Valid patient identification
- Initial clinical assessment
- Emergency cases receive immediate priority

2nd Chunk

Patient information is recorded in the Electronic Health Record (EHR) system.

3rd Chunk

Bed allocation follows the Bed Occupancy guidelines maintained by the Operations team.

Visually, the process looks like this:

Chunking breaks the structure of the data

When a clinician asks,

“What is the patient admission process?”

The vector database retrieves the chunks that seem most relevant, but the logical relationships between the admission policy, emergency triage, the EHR system, and bed allocation are lost. The model has to reconstruct them on every query. This isn’t a flaw in RAG. It remains one of the best techniques for searching large, unstructured collections like PDFs, research papers, support tickets, and historical records.

Curated organisational knowledge is different. Policies, procedures, APIs, and runbooks aren’t just text, they’re interconnected concepts. Rebuilding those links from fragmented chunks on every query adds needless complexity, and that is exactly the problem OKF was designed to solve.

What is the Open Knowledge Format (OKF)?

The ideas behind OKF did not originate with Google. Earlier in 2026, Andrej Karpathy introduced the concept of an LLM Wiki: instead of repeatedly retrieving raw documents, an AI agent maintains a curated knowledge base it can continuously read, update, and improve. His analogy caught on quickly in the AI community:

Obsidian is the IDE. The LLM is the programmer. The wiki is the codebase.

The idea is simple. Humans provide source material like documentation, policies, schemas, and runbooks, and the agent organises it into a structured wiki by writing summaries, connecting related concepts, and maintaining links. Those relationships become part of the knowledge base instead of being rediscovered on every query.

Google turned this community idea into an open specification. Rather than shipping another framework or SDK, it focused on standardising the knowledge itself. The result is OKF, a lightweight format that stores knowledge as ordinary Markdown files with minimal metadata and explicit links between concepts.

An OKF bundle is just a directory of Markdown documents, each representing one concept such as a policy, API, department, runbook, database table, or metric, connected through standard Markdown links. Unlike a vector database that infers relationships through embedding similarity, OKF preserves them explicitly, so an agent follows links rather than guessing.

Because everything is plain text, it fits existing developer workflows: version-controlled in Git, reviewed via pull requests, and searchable with standard tools. Next, we’ll build a bundle from scratch to see how it is organised.

Structure of an OKF Bundle

Now that we’ve understood the motivation behind OKF, let’s look at how an OKF bundle is actually organised. At its core, an OKF bundle is simply a directory of Markdown files. Each Markdown file represents one concept, such as a hospital policy, department, procedure, system, or operational metric. Every concept contains lightweight metadata followed by structured Markdown content. Related concepts are connected using standard Markdown links, allowing both humans and AI agents to navigate the knowledge base naturally.

The specification itself is intentionally minimal. It defines only a few conventions and avoids enforcing a rigid directory structure. This gives organisations the flexibility to organise knowledge in a way that best fits their domain while still producing bundles that can be understood by any OKF-compatible agent.

A typical OKF bundle contains the following components.

Component Purpose
index.md Serves as the primary entry point into the knowledge base. It provides an overview of the available concepts and helps agents navigate the bundle.
CHANGELOG.md (Optional) Records changes made to the knowledge base over time, making updates transparent and traceable.
Concept Files (.md) Each Markdown file represents a single concept such as a policy, procedure, API, department, metric, or system.
YAML Front Matter Stores metadata including the concept type, title, description, tags, ownership, and last updated timestamp.
Markdown Links Explicitly connect related concepts, transforming the knowledge base into a navigable graph instead of isolated documents.

A Typical OKF Folder Structure

Although the OKF specification does not mandate a particular directory layout, following a consistent folder hierarchy makes the knowledge base significantly easier to maintain and navigate. The same organisational principles apply regardless of the domain.

The following examples demonstrate how different organisations can structure their knowledge while following the same OKF conventions.

1: Hospital Knowledge Base

Hospital Knowlege Base

2: Software Engineering Knowledge Base

Software Engineer Knowledge Base

3: Manufacturing Knowledge Base

Manufacturing knowledge base

Although these examples belong to completely different industries, the underlying organisation remains remarkably similar. Every bundle begins with an index.md file that serves as the entry point, an optional CHANGELOG.md for tracking revisions, and a set of directories that group related concepts together.

This consistency is one of OKF’s biggest strengths. Once an AI agent understands how one OKF bundle is organised, it can navigate another bundle built using the same conventions with little or no additional adaptation.

Building an OKF Bundle

Now that we’ve explored the overall structure of an OKF bundle, let’s build one from scratch.

For the remainder of this article, we’ll use a fictional hospital called CityCare Hospital. Imagine we’re building an AI assistant that helps doctors, nurses, and hospital administrators answer operational questions. The assistant should understand admission policies, emergency procedures, hospital departments, internal systems, and operational metrics. Instead of storing this information inside a vector database, we’ll organise it as an OKF bundle.

We’ll begin by creating the root directory.

CityCare Root Directory

The index.md file acts as the entry point for both humans and AI agents.

# CityCare Hospital Knowledge Base

## Policies

- [Patient Admission Policy](policies/patient-admission.md)
- [Discharge Policy](policies/discharge-policy.md)

## Procedures

- [Emergency Triage](procedures/emergency-triage.md)
- [Blood Transfusion](procedures/blood-transfusion.md)

## Systems

- [Electronic Health Record](systems/ehr-system.md)

## Metrics

- [Bed Occupancy](metrics/bed-occupancy.md)

## Departments

- [Emergency Department](departments/emergency.md)

Rather than searching the entire repository, an AI agent can first read the index to understand what concepts exist before navigating to the relevant files. This simple design keeps the bundle organised while reducing unnecessary context during retrieval.

In the next section, we’ll create individual concept files and examine how YAML metadata, Markdown content, and links work together to make the knowledge base understandable for both humans and AI agents.

Creating an OKF Concept File

The building blocks of an OKF bundle are concept files. Each concept represents exactly one piece of knowledge, such as a policy, procedure, department, system, metric, or API. Keeping concepts focused makes them easier to maintain while allowing AI agents to retrieve only the information they need.

Every concept file consists of two parts:

  1. YAML Front Matter, which stores metadata about the concept.
  2. Markdown Content, which contains the actual knowledge along with links to related concepts.

Let’s create a concept file for the hospital’s patient admission policy.

---
type: policy
title: Patient Admission Policy
description: Guidelines for admitting patients into CityCare Hospital
tags:
- admissions
- patient-care
updated: 2026-06-15
---

# Patient Admission Policy

Patients arriving through the Emergency Department must complete an initial triage before admission.

## Admission Requirements

- Valid patient identification
- Initial clinical assessment completed
- Emergency cases receive immediate priority

## Related Concepts

- [Emergency Triage](../procedures/emergency-triage.md)
- [Electronic Health Record](../systems/ehr-system.md)

Notice that the file contains much more than plain text. The YAML section describes what kind of knowledge this file represents, while the Markdown body explains the concept in detail. Most importantly, the concept links to other related concepts inside the knowledge base. These links transform isolated documents into an interconnected knowledge graph that an AI agent can navigate.

Although OKF only requires the type field, adding additional metadata makes the bundle easier to organise and maintain.

Field Description
type Identifies the type of concept, such as policy, procedure, system, or metric. This is the only required field in the current specification.
title Human-readable title of the concept.
description Summary describing the concept.
tags Keywords that help organise related concepts.
updated Indicates when the concept was last modified.

Since the metadata is stored in YAML, both humans and AI agents can quickly understand what a document represents before reading its full content.

One of the biggest differences between OKF and traditional document storage is that concepts are explicitly connected using Markdown links.

For example, the admission policy references the emergency triage procedure and the Electronic Health Record (EHR) system.

Explicit Links Between Concepts

These relationships are intentionally created by the author. The agent does not have to infer them through semantic similarity because they already exist inside the knowledge base.

Another Example: Hospital Metric

Concept files are not limited to policies. The same structure can describe operational metrics, internal systems, departments, APIs, or runbooks.

Below is a concept describing the hospital’s Bed Occupancy Rate.

---
type: metric
title: Bed Occupancy Rate
description: Percentage of inpatient beds currently occupied
tags:
- operations
- hospital
updated: 2026-06-15
---

# Bed Occupancy Rate

The Bed Occupancy Rate measures the percentage of inpatient beds currently occupied.

## Formula
Occupied Beds / Total Available Beds × 100

## Data Source

Hospital Information System

## Owner

Operations Department

## Related Concepts

- [Emergency Department](../departments/emergency.md)
- [Patient Admission Policy](../policies/patient-admission.md)

Because every concept follows a consistent structure, an AI agent can quickly understand what the metric represents, how it is calculated, where the data comes from, and which other concepts are related to it.

How AI Agents Traverse an OKF Bundle

Once the knowledge base is organised into interconnected concept files, retrieval becomes much simpler than traditional document search.

Instead of searching thousands of document chunks, an AI agent follows a structured navigation process.

  1. Read the index.md file to understand the overall knowledge base.
  2. Identify the most relevant concept based on the user’s question.
  3. Open that concept file.
  4. Follow links to related concepts whenever additional context is needed.
  5. Generate the final response using only the relevant concepts.

The traversal process can be visualised as follows.

How a Agent Traverses OKF

Unlike a RAG pipeline, the agent does not begin by searching an embedding index. It starts from a curated entry point and progressively explores only the concepts that are relevant to the current task.

This approach preserves the relationships between concepts while keeping the amount of context sent to the language model relatively small.

Why This Works Well for AI Agents

The biggest advantage of OKF is that it allows developers to organise knowledge in the same way humans naturally think about it. A doctor reading the hospital’s documentation does not randomly jump between unrelated paragraphs. They begin with a policy, follow references to procedures, consult the relevant systems, and then arrive at the information they need. OKF enables AI agents to follow this same workflow.

Instead of reconstructing relationships from fragmented document chunks every time a question is asked, the agent navigates an explicit knowledge graph where those relationships have already been defined. This makes the retrieval process more deterministic, easier to audit, and significantly simpler to maintain.

Where RAG Still Excels

At this point, OKF sounds like an ideal solution for organising knowledge. It preserves relationships between concepts, keeps everything version-controlled, and allows AI agents to navigate curated documentation without relying on semantic search.

However, OKF has an important limitation. Someone has to curate every concept.

Every policy, procedure, system, metric, and department must be written, reviewed, and maintained. This works well for authoritative organisational knowledge, but it becomes impractical when the knowledge base grows to millions of documents.

Consider a hospital that has accumulated years of operational data.

  • Millions of Electronic Health Record (EHR) entries
  • Medical research papers
  • Clinical notes
  • Patient feedback
  • Incident reports
  • Internal emails
  • Meeting transcripts
  • Equipment maintenance logs

Organising every one of these documents into carefully curated OKF concept files would require an enormous amount of manual effort. Even if AI agents assisted with the curation process, much of this information changes continuously and is better suited for semantic search.

This is where Retrieval-Augmented Generation (RAG) continues to be the preferred solution.

RAG Solution

Rather than requiring documents to be manually organised, RAG indexes large collections of unstructured data using embeddings. When a question is asked, the system retrieves the most semantically relevant documents and provides them as context to the language model.

For example, consider the following questions:

  • Has anyone encountered this MRI scanner error before?
  • Find previous incident reports involving delayed laboratory results.
  • Summarise discussions about the new EHR rollout.
  • Search for all meeting notes mentioning patient transfer delays.

These questions cannot be answered from a small curated knowledge base. Instead, they require searching through thousands or even millions of documents where the answer could exist anywhere. This is exactly the type of problem RAG was designed to solve.

The strengths of each approach become much clearer when viewed side by side.

Feature OKF RAG
Best for Curated organisational knowledge Large collections of unstructured documents
Knowledge Source Markdown concept files Raw documents
Retrieval Deterministic navigation Semantic similarity search
Infrastructure File system + Git Embeddings + Vector Database
Version Control Native Git support Requires re-indexing after updates
Relationships Explicit links between concepts Inferred from retrieved chunks
Scalability Moderate Excellent
Explainability High Moderate

Neither approach is universally better than the other. They simply solve different problems. OKF provides structure and precision. RAG provides scale and flexibility. This naturally raises another question.

Do we really have to choose one over the other?

Fortunately, the answer is no.                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                             

Hybrid Knowledge Architecture: Combining OKF and RAG

In practice, the most effective AI systems use both OKF and RAG together.

Instead of treating them as competing technologies, modern agent architectures use each one where it performs best.

A simple way to think about this is the 80/20 principle.

  • The 20% of organisational knowledge that is critical, stable, and frequently referenced is stored as an OKF bundle.
  • The remaining 80% of large, unstructured information stays inside a traditional RAG pipeline.

This creates a layered knowledge architecture.

Hybrid Knowledge Architecture

The router determines which knowledge source is most appropriate for the incoming query.

Questions requiring authoritative and deterministic answers are routed to the OKF bundle.

Examples include:

  • What is the patient admission policy?
  • How is the Bed Occupancy Rate calculated?
  • Which system stores patient medical records?
  • What is the emergency blood transfusion procedure?

Each of these questions has a single authoritative answer maintained by the organisation.

On the other hand, exploratory questions are routed to the RAG pipeline.

For example:

  • Has anyone encountered this MRI scanner error before?
  • Find previous incident reports involving delayed laboratory results.
  • Summarise discussions about the EHR migration project.
  • Search meeting notes discussing patient discharge delays.

These questions require searching large collections of historical documents rather than consulting curated knowledge.

This hybrid architecture allows each system to focus on its strengths.

OKF Strengths RAG Strengths
Deterministic retrieval Semantic retrieval
Curated and authoritative knowledge Massive document collections
Explicit relationships between concepts Finds information using semantic similarity
Version-controlled with Git Continuously indexes new documents
Easy to audit and maintain Highly scalable

Perhaps the biggest advantage of this architecture is that the language model does not need to know where the information comes from. The agent simply requests the knowledge it needs.

The routing layer decides whether that knowledge should come from the OKF bundle or the vector database. Frameworks such as LangGraph, LangChain, or LlamaIndex make this routing straightforward by allowing developers to build workflows that choose the appropriate retrieval strategy based on the user’s query. As a result, AI agents gain the precision of curated knowledge without sacrificing the ability to search vast collections of unstructured information.

In other words, the future is not OKF versus RAG. It is OKF plus RAG, working together as complementary layers in a single knowledge architecture.

Conclusion

The Open Knowledge Format offers a simple, transparent way to organise knowledge for AI agents. By representing it as interconnected Markdown documents, OKF keeps organisational knowledge easy to understand, maintain, and version-control in Git. 

It suits curated information like policies, runbooks, and API docs, without replacing RAG, which still excels at semantic search across large, unstructured collections. 

Used together, the two cover far more ground than either alone. Ultimately, knowing where each approach fits is what lets developers build agents that are accurate, explainable, and production-ready.

Frequently Asked Questions

Q1. What is the Open Knowledge Format (OKF)?

A. It is an open specification Google introduced in June 2026 for how AI agents organise and exchange knowledge. An OKF bundle is a directory of Markdown files with lightweight YAML metadata and explicit links between concepts, so knowledge lives as plain text alongside your code rather than as vectors in a database.

Q2. How is OKF different from RAG?

A. RAG splits documents into chunks, embeds them, and retrieves the most semantically similar pieces at query time. OKF stores knowledge as interconnected concept files and lets an agent navigate by following author-defined links. RAG infers relationships; OKF keeps them explicit.

Q3. Does OKF replace RAG?

A. No. They solve different problems. OKF is best for curated, authoritative knowledge like policies, runbooks, and API docs. RAG is best for searching large, unstructured collections such as EHR entries, incident reports, and meeting notes. The article recommends using them together.

Data Scientist @ Analytics Vidhya | CSE AI and ML @ VIT Chennai
Passionate about AI and machine learning, I'm eager to dive into roles as an AI/ML Engineer or Data Scientist where I can make a real impact. With a knack for quick learning and a love for teamwork, I'm excited to bring innovative solutions and cutting-edge advancements to the table. My curiosity drives me to explore AI across various fields and take the initiative to delve into data engineering, ensuring I stay ahead and deliver impactful projects.