惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

T
The Blog of Author Tim Ferriss
Microsoft Azure Blog
Microsoft Azure Blog
S
SegmentFault 最新的问题
Schneier on Security
Schneier on Security
W
WeLiveSecurity
Webroot Blog
Webroot Blog
T
Threatpost
量子位
大猫的无限游戏
大猫的无限游戏
C
Cisco Blogs
腾讯CDC
N
News | PayPal Newsroom
T
Troy Hunt's Blog
T
Tailwind CSS Blog
Latest news
Latest news
K
KPMG report finds enterprise disconnect between AI and its ROI | CIO
The Register - Security
The Register - Security
Know Your Adversary
Know Your Adversary
CTFtime.org: upcoming CTF events
CTFtime.org: upcoming CTF events
Engineering at Meta
Engineering at Meta
SecWiki News
SecWiki News
MyScale Blog
MyScale Blog
GbyAI
GbyAI
Application and Cybersecurity Blog
Application and Cybersecurity Blog
A
Arctic Wolf
The GitHub Blog
The GitHub Blog
Threat Intelligence Blog | Flashpoint
Threat Intelligence Blog | Flashpoint
cs.CV updates on arXiv.org
cs.CV updates on arXiv.org
Help Net Security
Help Net Security
cs.AI updates on arXiv.org
cs.AI updates on arXiv.org
有赞技术团队
有赞技术团队
NISL@THU
NISL@THU
L
LINUX DO - 最新话题
雷峰网
雷峰网
P
Privacy International News Feed
Spread Privacy
Spread Privacy
Attack and Defense Labs
Attack and Defense Labs
N
News and Events Feed by Topic
月光博客
月光博客
V
V2EX
H
Help Net Security
博客园 - 三生石上(FineUI控件)
U
Unit 42
B
Blog
PCI Perspectives
PCI Perspectives
G
GRAHAM CLULEY
Stack Overflow Blog
Stack Overflow Blog
P
Proofpoint News Feed
T
Threat Research - Cisco Blogs
Last Week in AI
Last Week in AI

DEV Community

Authentication Security Deep Dive: From Brute Force to Salted Hashing (With Java Examples) Why AI Systems Don’t Fail — They Drift Spilling beans for how i learn for exam😁"Reinforcement Learning Cheat Sheet" I Replaced Chrome with Safari for AI Browser Automation. Here's What Broke (and What Finally Worked) How Python Borrows Other People's Work The $40 Architecture: Processing 1 Billion API Requests with 99.99% Uptime Vibe Coding: A Workflow Guide (From Zero to SaaS) Most webhook security guides protect the wrong side. The scary part is delivery. Headless CMS for TanStack Start: Build a Blog with Cosmic EU Age Verification App "Hacked in 2 Minutes" — What Actually Happened Comfy Cloud’s delete function does not actually remove files Running AI Models on GPU Cloud Servers: A Beginner Guide Event-driven media intelligence with AWS Step Functions and Bedrock I scored 500 AI prompts across 8 quality dimensions — here's what broke How to Call Google Gemini API from Next.js (Free Tier, No Backend Needed) The Portal Protocol: Reclaiming Human Connection in the Age of AI How to Fix Your Team's Scattered Knowledge Problem With a Self-Hosted Forum Intro to tc Cloud Functors: A Graph-First Mental Model for the Modern Cloud Designing Multi-Tenant Backends With Both Ownership and Team Access I Built a Neumorphic CSS Library with 77+ Components — Here's What I Learned PostgreSQL Performance Optimization: Why Connection Pooling Is Critical at Scale Cómo construí un SaaS multi-rubro para gestionar expensas en Argentina con FastAPI + Vue 3 🚀 I Built an Ethical Hacking Scanner Tool – Open Source Project I Replaced /usage and /context in Claude Code With a Single Statusline A Pythonic Way to Handle Emails (IMAP/SMTP) with Auto-Discovery and AI-Ready Design I Collected 8.9 Million Polymarket Price Points — Here's What I Found About How Markets Really Move EcoTrack AI — Carbon Footprint Tracker & Dashboard Everyone's Using AI. No One Agrees How. 5 self-hosted ebook managers worth trying in 2026 Building Your First AI Agent with LangChain: From Chatbot to Autonomous Assistant Common SOC 2 Failures (Real World) Stop Vibe-Checking Your AI App: A Practical Guide to Evals How to Use SonarQube and SonarScanner Locally to Level Up Your Code Quality Your Next To-Do App Is Dead — I Replaced Mine with an OpenClaw AI Sign a Nostr event in 60 lines of Python using coincurve — no nostr-sdk, no nbxplorer, no rust toolchain ITGC Audit Explained Like You’re in Big 4 Patch Tuesday abril 2026: Microsoft parcha 163 vulnerabilidades y un zero-day en SharePoint Stop scraping everything: a better way to track competitor price changes Listing on MCPize + the Official MCP Registry while routing payments OUTSIDE the marketplace — how I kept 100% of my x402 revenue Building an AI-Powered Risk Intelligence System Using Serverless Architecture Why We Ripped Function Overloading Out of Our AI Toolchain Testing AI-Generated Code: How to Actually Know If It Works SaaS Churn Is Killing Your Business. Here Is What to Do About It (Without a Support Team) The Speed of AI Is No Longer Linear - And Self-Improving Models Are Why How to Implement RBAC for MCP Tools: A Practical Guide for Engineering Teams From Standard Quote to Persuasive Proposal: AI Automation for Arborists I built a CLI that scaffolds complete multi-tenant SaaS apps Axios CVE-2025–62718: The Silent SSRF Bug That Could Be Hiding in Your Node.js App Right Now The dashboard that ended our friendship Data Pipelines Explained Simply (and How to Build Them with Python) The Hidden Cost of AI Systems Nobody Talks About. undefined vs undeclared, and how typeof behaves Switching from file-based jobs to NATS/Kafka in Rust without changing code io_uring Adventures: Rust Servers That Love Syscalls Why Agentic AI is Killing the Traditional Database The POUR principles of web accessibility for developers and designers Quantum Neural Network 3D — A Deep Dive into Interactive WebGL Visualization How To Install Caveman In Codex On macOS And Windows Automation Pipeline Reliability: Why Your Workflow Breaks When Nobody Is Watching I Built an 'Open World' AI Coding Agent — It Works From ANY Folder From Freelancing to Product: A Tech Service Company's SaaS Transformation China's AI Giants: Adding Tencent Hunyuan & ByteDance Doubao to AI University (74 Providers) On the Vibe Coders and Their Lies clerk: Auto-Summarize Your Claude Code Sessions AI Weekly — 2026/04/10–04/17 | The Model Lockdown Is Here, but the Toolchain Is the Real Battleground AI 週報 — 2026/04/10–2026/04/17 模型封鎖潮來了,但工具鏈才是真戰場 Maybe this is how Open-Source apps are born... 🚀 Fine-Tune LLMs with LoRA and QLoRA: 2026 Guide tRPC v11 + Next.js App Router: End-to-End Type Safety Without the Boilerplate ShadCN UI in 2026: Why I Stopped Installing Component Libraries and Started Owning My Components SaaS Billing in React Server Components: Stripe + Supabase Without a Single `useEffect` Join our DEV Weekend Challenge — $1,000 in Prizes Across TEN winners! Submissions Due April 20 at 6:59 AM UTC. Implementing FSRS Spaced Repetition in Flutter + Supabase — Adding Memory Science to an AI Learning App "I Texted My Localhost From the Train — Claude Code Fixed the Bug Before I Got Home" I Built a Sales Prep AI and It Went Deeper Than Expected Design to Code #2: One JSON, Eleven Outputs Solving the 100M-Row Problem: A Summary Table Pattern for High-Volume Push Notification Logs Flutter Web With Wasm: What Actually Changes For Developers I Built 50 Royalty-Free Soundtracks for My Side Project in a Weekend Using AI Music Generation The Vibe Coding Security Checklist: 7 Things to Check Before You Ship Stop Letting Googlebot Guess Fix Your React App's SEO Right Desconstruindo o Streaming do LinkedIn: Como Criar um Engine de Extração de Vídeo de Alta Performance com HLS e FFmpeg (EDA Part-1) EDA (Exploratory Data Analysis) Explained With Real Life — Why Looking at Your Data Is the Most Important Step in Machine Learning Brand Relationship Management at Scale: Our 4-Touch Outreach System for 200+ Brands Why String.fromEnvironment() Might Return an Empty String in Dart JGuardrails 1.0.0 — Hardening Java LLM Apps Against Jailbreaks, Toxicity, and Prompt Injection Plan and Schedule a Full Week of Threads Content From One Claude Conversation Coding Cat Oran Ep3, Five Tables Changed Everything Updated: BFF Pattern I'm done watching freelancers get buried by 200 proposals. So I'm building the alternative. This is my first post BFS Algorithm in Java Step by Step Tutorial with Examples Tracking LLM Pricing Monthly: An Open Dataset for 22 AI Models How We Measure Content ROI on a Comparison Site: Revenue Attribution Without Perfect Data Introducing Nova AI Ops: The AI-Native Operating System for SRE Teams I built a free desktop video downloader for Windows — Grabbit How Talkie OCR Helps Vision-Impaired & Dyslexic Users Read the World Around Them VRCFaceTracking安装和iPhone面捕配置教程,有bug Even CrowdStrike Can't See Your Agents The Automation Gold Rush: What n8n Workflows and Claude Are Opening Up for Developers Right Now
Beyond Keyword Search: How Synthadoc v0.2.0 Combines BM25 and Vector Search to Build a Smarter Domain Wiki
Paul Chen · 2026-04-27 · via DEV Community

What is Synthadoc?

Synthadoc is an open-source, LLM-powered wiki engine. Point it at your
organisation's documents - PDFs, PPTX, spreadsheets, DOCX, images, or web pages - and it builds a persistent, structured knowledge base your team can query, audit, and extend over time.

Unlike general-purpose RAG pipelines that retrieve raw chunks at query
time and discard results afterwards, Synthadoc compiles knowledge at
ingest time into a living wiki that grows smarter and more consistent
with every new source. The core lifecycle is:

  • Ingest: extract and synthesise facts from any source format (PDF, XLSX, PNG, web URL)
  • Detect: flag contradictions with existing pages and quarantine them for review
  • Link: connect related pages and surface knowledge gaps
  • Query: answer questions with hybrid BM25 + optional vector search, citing the pages used
  • Lint: resolve contradictions and surface orphan pages for human or automated action

Synthadoc is designed for organisations that need domain-specific, auditable knowledge management: legal teams tracking regulatory
precedent, financial analysts maintaining market research, engineering
groups documenting system behaviour, and research teams building
institutional memory that persists beyond individual contributors.

Synthadoc v0.2.0 is released last week, it scales seamlessly while maintaining accuracy through autonomous self-optimization.

👉 Synthadoc GitHub: https://github.com/axoviq-ai/synthadoc

Introduction

Most LLM knowledge tools take one of two approaches to retrieval: pure
keyword search (fast but vocabulary-dependent) or pure vector/semantic
search (flexible but resource-intensive). In practice, both have
meaningful blind spots.

Synthadoc v0.2.0 ships a hybrid retrieval pipeline that uses BM25 as a
fast, precise first-pass filter and optional vector re-ranking as a
semantic second pass. The result is a system that is accurate on
exact-match queries, robust on paraphrased or conceptual queries, and
fast enough to run on a laptop with no cloud dependency.

This post explains how each technique works, where each one falls short
alone, why the hybrid matters for a persistent domain wiki, and how
Synthadoc v0.2.0 layers query decomposition and knowledge gap detection
on top.

What Is BM25?

BM25 (Best Match 25) is a probabilistic ranking function. It scores a
page relative to a query by counting how often query terms appear in the
page, discounting terms that appear in almost every page, and penalising
very long pages for artificially inflated counts. BM25 is the retrieval
backbone of Elasticsearch, Lucene, and most production search systems.

Scoring intuition

Where BM25 falls short

  • Vocabulary mismatch: query says "contributions", page says "pioneered". Score: near zero.
  • Synonyms: "ML" and "machine learning" are different tokens.
  • Conceptual distance: "reasoning under uncertainty" and "probabilistic inference" are semantically identical but lexically distant.

On a domain wiki ingested from diverse sources (papers, docs, blog posts, PDFs), the same concept will be described in many different vocabularies. BM25 alone misses a meaningful fraction of relevant pages.

What Is Vector Search?

Vector (semantic) search encodes text into dense numerical embeddings using a neural language model. Semantically similar texts land close together in that high-dimensional space regardless of surface wording.
Similarity is measured as cosine distance between the query vector and each page vector.

Embedding intuition

The two sentences share almost no keywords, but their vectors point in nearly the same direction because the model understands they describe the same concept.

Where vector search falls short alone

  • Cold-start penalty: embedding thousands of pages takes time and compute; BM25 is instant.
  • Exact-match dilution: specific product names or identifiers can be blurred by semantic proximity.
  • Domain drift: general-purpose models may not distinguish highly specific domain terminology.
  • Resource requirement: needs a model (~130 MB for bge-small-en-v1.5) and inference at query time.

BM25 vs. Vector Search: Side-by-Side

Dimension BM25 Vector / Semantic
Matching strategy Matching strategy Exact term overlap (TF x IDF) Semantic similarity (cosine distance)
Vocabulary required Query words must appear in page Paraphrases and synonyms handled
Speed Microseconds -- no model needed Milliseconds -- model inference required
Setup cost Zero -- pure algorithm ~130 MB model download (one-time)
Exact-match queries Excellent Good
Synonym / paraphrase Often misses Handles well
Domain terminology Good if terms match Depends on model training
Interpretability Score is explainable Black-box similarity
Best for Known vocabulary, structured content Conceptual queries, diverse sources

How Synthadoc Combines Both

Synthadoc uses a hybrid pipeline where BM25 and vector search are not alternatives - they are sequential layers. BM25 does the heavy filtering; vector re-ranks the survivors.

The retrieval pipeline

Query decomposition: why it matters for retrieval

Before any search happens, Synthadoc v0.2.0 breaks compound questions into focused sub-questions via an LLM call. Each sub-question runs its own BM25 (and vector) search in parallel. Results are merged by best score per page before synthesis. One complex query can retrieve from multiple distinct parts of the wiki simultaneously.

Example: Query: "Compare Turing's contributions with Von Neumann's
architecture"
-> Decomposed: ["Turing contributions computing"] | ["Von Neumann
architecture design"]
-> Two parallel BM25 searches -> merged candidates -> one synthesised
answer

Knowledge gap detection

After retrieval, Synthadoc evaluates three independent signals:

  • Fewer than 3 pages retrieved - the wiki barely covers the topic
  • Max BM25 score below configurable threshold (default: 2.0) - weak keyword overlap
  • Fewer than 2 candidates contain key nouns from the question - off-topic matches

When a gap fires, Synthadoc generates targeted web search suggestions and surfaces them as an Obsidian callout or CLI tip, creating a feedback loop that makes the wiki progressively denser over time.

Practical Examples in Synthadoc

Example 1: BM25 exact match (no vector needed)

Wiki page: "Alan Turing - Enigma and the Bombe Machine"

Query: "Bombe machine Enigma decryption"

BM25 succeeds: BM25 score: HIGH - "Bombe", "machine", "Enigma",
"decryption" all present.
Result: page retrieved correctly. Vector re-ranking not required.`**

Example 2: BM25 misses, vector rescues

Wiki page: "Alan Turing - Theoretical Foundations of Modern Computers"

Query: "What were Turing's contributions to computing?"

BM25 misses, Vector rescues: BM25 score: LOW - "contributions" and
"computing" absent from the page.
Vector cosine score: HIGH - embeddings are semantically close.
Result: page retrieved correctly after re-ranking. BM25 alone would have > missed it.

Example 3: Knowledge gap fires, ingest suggestion generated

Wiki: finance domain. Query: "What is the impact of quantitative easing on inflation?"

Gap detected: BM25: 1 page returned, score 0.8 (below threshold 2.0)
Knowledge gap detected. Synthadoc generates:
synthadoc ingest "search for: quantitative easing inflation impact" -w finance-wiki
synthadoc ingest "search for: central bank monetary policy effects" -w finance-wiki
After ingest and re-query: 7 pages returned, fully synthesised answer.

Enabling Vector Search in Synthadoc

BM25 is the default - zero setup, zero dependencies. To add vector re-ranking:

Step 1: Install fastembed

pip install fastembed

Step 2: Enable in config

[search]
vector = true
vector_top_candidates = 20 # BM25 pool size before re-ranking

Step 3: Restart the server

synthadoc serve -w my-wiki

On first start, Synthadoc downloads BAAI/bge-small-en-v1.5 (~130 MB) once and embeds existing pages in the background. BM25 stays active throughout - no downtime. If the model is unavailable, the system falls back to BM25 silently.

Synthadoc v0.2.0: Full Feature Summary

Feature What it does
Query decomposition Compound questions split into parallel BM25 sub-queries, merged before synthesis
Vector re-ranking Opt-in semantic re-ranking (BAAI/bge-small-en-v1.5 via fastembed)
Knowledge gap detection 3-signal gap check; auto-generates targeted ingest suggestions as Obsidian callout
Web search decomposition Broad search topics split into focused Tavily queries; URL deduplication and cap
Per-model cost tracking Per-token rate table; ingest + query cost in audit.db, CLI, and Obsidian
Query audit trail Full query history with sub-question count, tokens, cost, timestamp
Obsidian live web search view Real-time polling panel: phase, pages created, URL errors as fan-out completes
8 new Obsidian commands 15 commands total: lint, auto-resolve, job retry/purge, audit history, scaffold
MiniMax support M2.5/M2.7 reasoning models with reasoning_content fallback for structured output
Rate-limit requeue 429 responses requeue job (retry budget preserved); fail-fast on daily quota
Job crash recovery in_progress jobs at shutdown auto-reset to pending on next startup
Bulk job cancel Cancel all pending jobs in one operation via CLI or API

How Synthadoc Compares to Alternatives

Most LLM knowledge tools are general-purpose RAG pipelines that retrieve raw chunks at query time with no persistent synthesis. Synthadoc compiles knowledge at ingest time, maintains a living wiki, and is designed for domain-specific, auditable deployments.

Capability Synthadoc v0.2.0 LlamaIndex / LangChain Notion AI Obsidian Copilot
Ingest-time synthesis Compiled wiki Raw chunks at query time Page-level only None
Domain scope filtering purpose.md Manual None None
Multi-model support 6 providers Many providers OpenAI only OpenAI / Ollama
Audit trail Full SQLite audit None built-in None None
Cost tracking Per-token, per-op Manual / callback Opaque None
Offline / local Fully local Depends on provider Cloud only Ollama
Obsidian-native output Wikilinks, Dataview None Notion-only Read-only
HTTP API + MCP server Built-in Manual wiring Proprietary API None
Contradiction detection Automated None None None
Query decomposition Parallel BM25 Manual chains None None
Knowledge gap detection Auto-suggestions None None None
Extensible skills Drop-in folders Custom loaders None None
Licence AGPL-3.0 open source Apache-2.0 Proprietary SaaS MIT

Enterprise and Domain-Specific Readiness

Synthadoc is built for organisations that need a knowledge system they control, audit, and deploy into existing infrastructure - not a SaaS black box.

Concrete use cases:

  • Legal - track regulatory updates, case precedents, and compliance requirements across jurisdictions. New ruling ingested, old page flagged contradicted, compliance team reviews.
  • Finance - build a living market research wiki from analyst reports, earnings calls, and regulatory filings. Query with natural language, get cited answers with full audit trail.
  • Engineering - maintain a persistent runbook that absorbs incident post-mortems, architecture decision records, and API docs. Contradiction detection prevents stale documentation from accumulating.
  • Research - aggregate papers, datasets, and notes into a structured knowledge base. Knowledge gap detection surfaces what the team does not yet know and generates targeted ingest suggestions.

Domain specificity

Every wiki defines its own scope via purpose.md. The LLM reads this before every ingest decision and rejects out-of-scope sources cleanly. A legal wiki does not absorb marketing copy. A financial wiki does not absorb engineering runbooks.

Auditability

Every ingest, query, contradiction detection, and auto-resolution is written to an append-only SQLite audit trail with token counts, cost, timestamps, and page-level actions -- all queryable from the CLI or Obsidian audit commands.

synthadoc audit history -w my-wiki # ingest records

synthadoc audit cost -w my-wiki # token spend breakdown

synthadoc audit events -w my-wiki # contradiction, gate, resolution
events

Product and grid readiness

Synthadoc exposes the same operations across four surfaces sharing a single agent and storage layer:

  • CLI: for operators, automation scripts, and CI pipelines
  • HTTP REST API: for product integrations and custom front-ends
  • MCP server: for direct agent-to-agent communication
  • Obsidian plugin: for knowledge workers doing active research

Hook scripts fire on lifecycle events (on_ingest_complete, on_lint_complete), enabling event-driven automation: post a Slack summary when ingest completes, trigger a downstream build when a key page changes, or chain into a broader orchestration pipeline. Cron scheduling is built in, and multi-wiki isolation means each team or domain runs on its own port with its own audit trail.

Synthadoc in Agentic Autonomous Systems

Synthadoc is purpose-built to serve as the persistent knowledge layer for LLM agent systems. Where an agent's context window is ephemeral and limited, Synthadoc's wiki is persistent, structured, and queryable -it gives agents a long-term memory that survives across sessions, scales to millions of tokens of accumulated knowledge, and is fully auditable.

Agent integration via MCP

The built-in MCP (Model Context Protocol) server exposes ingest, query, and lint as native tool calls. An agent running in any MCP-compatible host - Claude, GPT-4o, a custom LangChain pipeline - can:

  • Call query to retrieve cited, synthesised answers from accumulated knowledge before acting
  • Call ingest to push new findings, research results, or external documents back into the wiki
  • Call lint to check for contradictions introduced by new data before committing to a decision

Event-driven agent pipelines

Hook scripts fire on lifecycle events and can trigger downstream agent actions:

  • on_ingest_complete: a downstream agent reads newly created pages and decides whether to trigger follow-up ingests or alert a human reviewer
  • on_lint_complete: an orchestrator agent receives contradiction and orphan reports and routes resolution tasks to specialised sub-agents

Example pipeline: a web-crawling agent ingests raw URLs; Synthadoc synthesises and deduplicates; a reporting agent queries the updated wiki and posts a daily briefing - all without human intervention.

Persistent domain memory for multi-agent systems

In multi-agent architectures, shared knowledge is a coordination bottleneck. Synthadoc solves this by acting as a shared, structured memory store:

  • Multiple agents read from the same wiki via parallel HTTP queries - no shared state management required
  • One agent's ingest results are immediately available to all agents querying the same wiki
  • Multi-wiki isolation means separate agent clusters maintain scoped knowledge without interference
  • The audit trail provides a complete record of which agent ingested what, when, at what cost - making multi-agent systems auditable by design

Because Synthadoc is self-hosted and open-source, teams building autonomous systems retain full control over data residency, model selection, and cost - a critical requirement for enterprise agentic
deployments.

Try Synthadoc v0.2.0

Synthadoc v0.2.0 is available now on GitHub under the AGPL-3.0 licence. BM25 search works out of the box. Vector re-ranking is one pip install away. The Gemini free tier means you can run a full ingest-and-query cycle at zero cost.

Feedback welcome: Feedback, issues, and contributions are very welcome. Open an issue on GitHub or start a discussion - the roadmap is shaped by what users need.

👉 README: https://github.com/axoviq-ai/synthadoc#readme
👉 Quick-start guide: https://github.com/axoviq-ai/synthadoc/blob/main/docs/user-quick-start-guide.md
👉 Design document: https://github.com/axoviq-ai/synthadoc//blob/main/docs/design.md
👉 Release notes: https://github.com/axoviq-ai/synthadoc/releases/tag/v0.2.0