惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

Hacker News: Ask HN
Hacker News: Ask HN
D
DataBreaches.Net
Microsoft Security Blog
Microsoft Security Blog
U
Unit 42
V
Visual Studio Blog
GbyAI
GbyAI
云风的 BLOG
云风的 BLOG
博客园 - Franky
C
CXSECURITY Database RSS Feed - CXSecurity.com
大猫的无限游戏
大猫的无限游戏
P
Privacy & Cybersecurity Law Blog
T
The Exploit Database - CXSecurity.com
Simon Willison's Weblog
Simon Willison's Weblog
L
LangChain Blog
I
Intezer
V2EX - 技术
V2EX - 技术
Google DeepMind News
Google DeepMind News
T
Threat Research - Cisco Blogs
Apple Machine Learning Research
Apple Machine Learning Research
V
V2EX
腾讯CDC
博客园 - 【当耐特】
Know Your Adversary
Know Your Adversary
TaoSecurity Blog
TaoSecurity Blog
cs.AI updates on arXiv.org
cs.AI updates on arXiv.org
F
Fortinet All Blogs
Project Zero
Project Zero
Blog — PlanetScale
Blog — PlanetScale
S
Security @ Cisco Blogs
量子位
M
MIT News - Artificial intelligence
美团技术团队
C
Cisco Blogs
S
Schneier on Security
Recent Commits to openclaw:main
Recent Commits to openclaw:main
G
Google Developers Blog
N
News and Events Feed by Topic
MongoDB | Blog
MongoDB | Blog
The Hacker News
The Hacker News
H
Help Net Security
S
Secure Thoughts
Scott Helme
Scott Helme
SecWiki News
SecWiki News
T
Troy Hunt's Blog
cs.CV updates on arXiv.org
cs.CV updates on arXiv.org
博客园 - 叶小钗
O
OpenAI News
Application and Cybersecurity Blog
Application and Cybersecurity Blog
博客园 - 司徒正美
T
Tenable Blog

DEV Community

Authentication Security Deep Dive: From Brute Force to Salted Hashing (With Java Examples) Why AI Systems Don’t Fail — They Drift Spilling beans for how i learn for exam😁"Reinforcement Learning Cheat Sheet" I Replaced Chrome with Safari for AI Browser Automation. Here's What Broke (and What Finally Worked) How Python Borrows Other People's Work The $40 Architecture: Processing 1 Billion API Requests with 99.99% Uptime Vibe Coding: A Workflow Guide (From Zero to SaaS) Most webhook security guides protect the wrong side. The scary part is delivery. Headless CMS for TanStack Start: Build a Blog with Cosmic EU Age Verification App "Hacked in 2 Minutes" — What Actually Happened Comfy Cloud’s delete function does not actually remove files Running AI Models on GPU Cloud Servers: A Beginner Guide Event-driven media intelligence with AWS Step Functions and Bedrock I scored 500 AI prompts across 8 quality dimensions — here's what broke How to Call Google Gemini API from Next.js (Free Tier, No Backend Needed) The Portal Protocol: Reclaiming Human Connection in the Age of AI How to Fix Your Team's Scattered Knowledge Problem With a Self-Hosted Forum Intro to tc Cloud Functors: A Graph-First Mental Model for the Modern Cloud Designing Multi-Tenant Backends With Both Ownership and Team Access I Built a Neumorphic CSS Library with 77+ Components — Here's What I Learned PostgreSQL Performance Optimization: Why Connection Pooling Is Critical at Scale Cómo construí un SaaS multi-rubro para gestionar expensas en Argentina con FastAPI + Vue 3 🚀 I Built an Ethical Hacking Scanner Tool – Open Source Project I Replaced /usage and /context in Claude Code With a Single Statusline A Pythonic Way to Handle Emails (IMAP/SMTP) with Auto-Discovery and AI-Ready Design I Collected 8.9 Million Polymarket Price Points — Here's What I Found About How Markets Really Move EcoTrack AI — Carbon Footprint Tracker & Dashboard Everyone's Using AI. No One Agrees How. 5 self-hosted ebook managers worth trying in 2026 Building Your First AI Agent with LangChain: From Chatbot to Autonomous Assistant Common SOC 2 Failures (Real World) Stop Vibe-Checking Your AI App: A Practical Guide to Evals How to Use SonarQube and SonarScanner Locally to Level Up Your Code Quality Your Next To-Do App Is Dead — I Replaced Mine with an OpenClaw AI Sign a Nostr event in 60 lines of Python using coincurve — no nostr-sdk, no nbxplorer, no rust toolchain ITGC Audit Explained Like You’re in Big 4 Patch Tuesday abril 2026: Microsoft parcha 163 vulnerabilidades y un zero-day en SharePoint Stop scraping everything: a better way to track competitor price changes Listing on MCPize + the Official MCP Registry while routing payments OUTSIDE the marketplace — how I kept 100% of my x402 revenue Building an AI-Powered Risk Intelligence System Using Serverless Architecture Why We Ripped Function Overloading Out of Our AI Toolchain Testing AI-Generated Code: How to Actually Know If It Works SaaS Churn Is Killing Your Business. Here Is What to Do About It (Without a Support Team) The Speed of AI Is No Longer Linear - And Self-Improving Models Are Why How to Implement RBAC for MCP Tools: A Practical Guide for Engineering Teams From Standard Quote to Persuasive Proposal: AI Automation for Arborists I built a CLI that scaffolds complete multi-tenant SaaS apps Axios CVE-2025–62718: The Silent SSRF Bug That Could Be Hiding in Your Node.js App Right Now The dashboard that ended our friendship Data Pipelines Explained Simply (and How to Build Them with Python) The Hidden Cost of AI Systems Nobody Talks About. undefined vs undeclared, and how typeof behaves Switching from file-based jobs to NATS/Kafka in Rust without changing code io_uring Adventures: Rust Servers That Love Syscalls Why Agentic AI is Killing the Traditional Database The POUR principles of web accessibility for developers and designers Quantum Neural Network 3D — A Deep Dive into Interactive WebGL Visualization How To Install Caveman In Codex On macOS And Windows Automation Pipeline Reliability: Why Your Workflow Breaks When Nobody Is Watching I Built an 'Open World' AI Coding Agent — It Works From ANY Folder From Freelancing to Product: A Tech Service Company's SaaS Transformation China's AI Giants: Adding Tencent Hunyuan & ByteDance Doubao to AI University (74 Providers) On the Vibe Coders and Their Lies clerk: Auto-Summarize Your Claude Code Sessions AI Weekly — 2026/04/10–04/17 | The Model Lockdown Is Here, but the Toolchain Is the Real Battleground AI 週報 — 2026/04/10–2026/04/17 模型封鎖潮來了,但工具鏈才是真戰場 Maybe this is how Open-Source apps are born... 🚀 Fine-Tune LLMs with LoRA and QLoRA: 2026 Guide tRPC v11 + Next.js App Router: End-to-End Type Safety Without the Boilerplate ShadCN UI in 2026: Why I Stopped Installing Component Libraries and Started Owning My Components SaaS Billing in React Server Components: Stripe + Supabase Without a Single `useEffect` Join our DEV Weekend Challenge — $1,000 in Prizes Across TEN winners! Submissions Due April 20 at 6:59 AM UTC. Implementing FSRS Spaced Repetition in Flutter + Supabase — Adding Memory Science to an AI Learning App "I Texted My Localhost From the Train — Claude Code Fixed the Bug Before I Got Home" I Built a Sales Prep AI and It Went Deeper Than Expected Design to Code #2: One JSON, Eleven Outputs Solving the 100M-Row Problem: A Summary Table Pattern for High-Volume Push Notification Logs Flutter Web With Wasm: What Actually Changes For Developers I Built 50 Royalty-Free Soundtracks for My Side Project in a Weekend Using AI Music Generation The Vibe Coding Security Checklist: 7 Things to Check Before You Ship Stop Letting Googlebot Guess Fix Your React App's SEO Right Desconstruindo o Streaming do LinkedIn: Como Criar um Engine de Extração de Vídeo de Alta Performance com HLS e FFmpeg (EDA Part-1) EDA (Exploratory Data Analysis) Explained With Real Life — Why Looking at Your Data Is the Most Important Step in Machine Learning Brand Relationship Management at Scale: Our 4-Touch Outreach System for 200+ Brands Why String.fromEnvironment() Might Return an Empty String in Dart JGuardrails 1.0.0 — Hardening Java LLM Apps Against Jailbreaks, Toxicity, and Prompt Injection Plan and Schedule a Full Week of Threads Content From One Claude Conversation Coding Cat Oran Ep3, Five Tables Changed Everything Updated: BFF Pattern I'm done watching freelancers get buried by 200 proposals. So I'm building the alternative. This is my first post BFS Algorithm in Java Step by Step Tutorial with Examples Tracking LLM Pricing Monthly: An Open Dataset for 22 AI Models How We Measure Content ROI on a Comparison Site: Revenue Attribution Without Perfect Data Introducing Nova AI Ops: The AI-Native Operating System for SRE Teams I built a free desktop video downloader for Windows — Grabbit How Talkie OCR Helps Vision-Impaired & Dyslexic Users Read the World Around Them VRCFaceTracking安装和iPhone面捕配置教程,有bug Even CrowdStrike Can't See Your Agents The Automation Gold Rush: What n8n Workflows and Claude Are Opening Up for Developers Right Now
Why I built rtfstruct, fifteen years after writing the RTF parser inside Scrivener
Lee Powell · 2026-05-05 · via DEV Community

Lee Powell · Architect of Scrivener and Scapple · Lumen & Lever

Most AI document pipelines fail before the model is ever called. Tables become paragraphs. Lists collapse into prose. Annotations are detached from context. Page references disappear. Source traceability is replaced by a confidence score. The structure that gave the document its meaning is gone before retrieval runs, and no retrieval recovers it.

This is the layer that gets underestimated. I have worked on it for a long time. Long before retrieval-augmented generation existed, I wrote the production C++ RTF reader and writer that ships inside Scrivener for Windows and Linux, used by hundreds of thousands of long-form writers across novels, dissertations, and screenplays. That code was eventually sold as a white-label engagement to Literature & Latte, who continue to maintain Scrivener today.

RTF is not a glamorous format. It is also not going away. It is still the wire format inside Microsoft Outlook for rich-text email. It is still produced by court reporting systems, medical records platforms, government archives, and twenty-year-old legal practice management systems. When a law firm pulls a thousand contracts out of an old document store, a meaningful portion of them are RTF. When a hospital exports decades of clinical notes, a meaningful portion are RTF. When you scrape an Outlook MSG file, the body is RTF.

The standard pipeline path converts these documents to plain text immediately. The structure goes. Tables become paragraphs. Section headings become bold lines indistinguishable from body emphasis. Numbered clauses lose their numbering. Footnotes lose their links to the text they annotate. The model performs less well, the answers are less reliable, and the cause sits a layer below where anyone is looking.

I built rtfstruct to fix that layer. It is a Python 3.11+ RTF reader and writer that produces a neutral document AST, preserves structure all the way through, exposes diagnostics rather than swallowing them, and supports clean roundtrip back to RTF. Apache-2.0. Part of Sourcetrace by Lumen & Lever, the document structure layer for AI pipelines that I now run as a consultancy.

What follows is why the layer matters, what is wrong with the existing tools, what rtfstruct does differently, and why ten years of writing production RTF code for a writing application turned out to be the right preparation for an AI ingestion problem.

Why structure is not optional

The phrase that captures the architectural mistake is "structure-before-model." When a workflow involves structurally rich documents, the first decision is not which model. It is what intermediate representation. A blood test report is not text. It is a structured clinical record with analytes, values, units, reference ranges, abnormal flags, methods, and trend history. The same is true of leases, bank statements, invoices, pathology reports, contracts. Each of them carries its meaning in the structure. Flatten the structure and the meaning becomes inferred rather than read.

The pipeline then asks the model to reconstruct probabilistically what the document already contained as deterministic structure. The result is silent error. Reference ranges honoured in one record and missed in another. Unit conversions correct for SI units and silently wrong for imperial. Clause cross-references followed in clean documents and lost in legacy ones. Nothing in the model's output makes the failure visible. The diagnostic surface was the source structure, and the source structure was discarded at ingestion.

The model should reason over the AST, not over the document. Where the source has structure, structure is the system. The model is the consultant the system calls when structure alone cannot answer the question.

That is the doctrine. It is also the reason a tool like rtfstruct exists at all. Flatten RTF to text before AI sees it and the model's job is harder, the evaluation surface is smaller, and the audit trail is unfit for production. Preserve structure into an AST and the same workflow ships, evaluates, and audits cleanly.

What the existing Python tools actually do

There are a handful of Python libraries in the RTF space. None of them does what an AI ingestion pipeline needs. The honest landscape:

Library What it does Gap
striprtf Strips RTF to plain text. Lightweight, popular, useful for quick conversion. Discards all structure. By design.
PyRTF / pyrtf-ng Generates RTF from Python. Writer-only. Pyrtf-ng is largely abandoned. Cannot read RTF at all.
rtfparse Decapsulates HTML embedded inside RTF (mainly for Outlook MSG bodies). Specialised for one use case. Not a general parser.
oletools rtfobj Extracts embedded objects from RTF for malware analysis. Forensic tool, not a document parser.
Aspose.Words Commercial Python wrapper around .NET. Handles RTF among many formats. Commercial license. .NET runtime dependency. Closed source.
rtfparserkit Solid RTF parser, listener-based. Java only. Not Python.

The gap has sat in the Python ecosystem for years. There is no AST-first RTF reader and writer designed for structured pipelines, with first-class diagnostics and source spans, that is also open source. The closest thing is striprtf, which is excellent at exactly the opposite of what AI ingestion needs.

That is the gap rtfstruct fills.

What rtfstruct does differently

Four things matter. None of them are individually exotic. The combination is the point.

1. The AST is the public contract

rtfstruct parses RTF into a neutral document AST that preserves paragraphs, inline styles, lists, tables, links, fields, footnotes, endnotes, annotations, images, metadata, source spans, and recoverable diagnostics. Every other operation in the library is defined against the AST. JSON export, Markdown export, RTF roundtrip, and integration helpers all read from the AST. It is not an internal representation that gets discarded after parsing. It is the artefact.

from rtfstruct import parse_rtf

document = parse_rtf(r"{\rtf1\ansi Hello, \b world\b0!}")

print(document.to_json())
print(document.to_markdown())
print(document.to_rtf())

Enter fullscreen mode Exit fullscreen mode

The AST distinguishes between a heading paragraph and a body paragraph, between a list item and a regular paragraph, between a footnote reference and a footnote body, between a table cell and a table row. None of that is in the rendered text. All of it is in the source RTF. A pipeline that sees the AST sees the document. A pipeline that sees flattened text sees a wall of words.

2. Diagnostics are returned with the document

RTF in production is messy. Twenty-year-old legacy documents have malformed control words, broken Unicode escapes, codepage mismatches. Most parsers either fail loudly or silently drop the affected content. Neither is what a production pipeline needs.

rtfstruct returns diagnostics as part of the document object. If a malformed Unicode escape is recovered, the recovered character comes back along with a diagnostic carrying the severity, code, message, and source location. The pipeline then makes explicit decisions: log it, surface it for human review, reject the document, or proceed with confidence flagged.

from rtfstruct import parse_rtf

document = parse_rtf(r"{\rtf1\u999999?}")

for diagnostic in document.diagnostics:
    print(diagnostic.severity.value, diagnostic.code, diagnostic.message)

Enter fullscreen mode Exit fullscreen mode

The value of this is invisible until the production system encounters its first malformed document. After that it becomes the difference between a pipeline that fails opaquely and a pipeline that fails informatively.

3. Source spans map AST nodes back to byte offsets

For tools that need to highlight a region in the original RTF (legal review interfaces, document comparison tools, evidence-traceable AI systems), source spans are mandatory. rtfstruct supports them as an opt-in parser option. When enabled, every AST node carries a span pointing to the byte range in the source RTF that produced it. This is the foundation for the kind of source traceability that production AI systems need but rarely build, because retrofitting it later is structurally impossible.

from rtfstruct import ParserOptions, parse_rtf

document = parse_rtf(
    r"{\rtf1 Hello}",
    options=ParserOptions(track_spans=True),
)

Enter fullscreen mode Exit fullscreen mode

4. Roundtrip without semantic loss

The reader and the writer share the same AST. A document parsed in, edited in place, and written back out preserves the structural choices it carried. This sounds straightforward and it is not. Most parsers that also write tend to lose information on the round trip. Inline style runs collapse, table cell properties drift, list numbering restarts. rtfstruct is tested for semantic roundtrip across inline formatting, metadata, fields and links, footnotes, annotations, lists, tables, images, and Unicode recovery.

from rtfstruct import read_rtf, write_rtf

document = read_rtf("input.rtf")
# inspect, modify, validate, classify...
write_rtf(document, "output.rtf")

Enter fullscreen mode Exit fullscreen mode

Why ten years of Scrivener was the right preparation

Maintaining a parser in production for a long time produces a particular kind of engineering scar tissue. RTF is a 38-year-old format with hundreds of control words, dozens of edge cases that only appear in real documents, and at least three different lineages of generators producing subtly different output (Microsoft Word, OpenOffice, and various enterprise systems built up over decades). The official specification documents some of this. The rest is learned by debugging support tickets from a writer in Iceland whose decade-old document refuses to load correctly.

The Scrivener parser went through that learning the hard way. Pressure-tested across hundreds of thousands of writers, on Windows and Linux, on documents ranging from short stories to thousand-page novels, dissertations with mixed-language sections, screenplays with industry-specific formatting, academic papers with footnotes and citations and embedded equations. By the time it shipped to production it handled malformed documents, codepage drift, Unicode escapes outside valid ranges, and recovery from errors that simpler parsers would treat as fatal. That code was eventually sold as a white-label engagement to Literature & Latte.

rtfstruct is not a port of the Scrivener parser. It is a fresh codebase, written for Python, designed for structured AI pipelines rather than for an interactive writing application. The thinking behind it is shaped by the ten years I spent watching writers feed Scrivener documents that should not have parsed and watching the parser handle them anyway. The decisions about what to recover, what to flag, what to expose as diagnostics, and what shape the AST should take are decisions made before, in production, under load, with real users. That experience compresses years of design discovery into a starting point most parser projects do not have.

A note on provenance: I do not maintain Scrivener and have not seen its codebase in years. The intellectual property was sold to Literature & Latte. The lessons stayed with me. rtfstruct is a different library, for a different purpose, written for a different language, drawing on the same engineering instinct that produced the original.

What this is for, and what it is not

rtfstruct is for systems where document structure still matters. AI ingestion pipelines, RAG systems, legal discovery, banking and financial archives, forensic document tracing, publishing pipelines, long-form document intelligence, and any pipeline that processes RTF as input rather than discarding it as a legacy format.

It is not a plain-text stripper. If all you need is the words, striprtf is excellent and you should use it. It is not a Markdown converter for casual use. If your input is well-formed contemporary RTF and your output is human-readable Markdown, several lighter tools will do the job. It is not a renderer. There is no HTML output, no styled rendering, no display library. It is an AST reader and writer.

The library is at version 0.1, currently labelled pre-alpha because the API will continue to evolve as integration patterns surface. The core reader, AST, JSON exporter, Markdown exporter, and RTF writer are working today. Tests cover inline formatting, metadata, fields and links, footnotes, annotations, lists, tables, images, Unicode and codepage recovery, diagnostics, source spans, and semantic roundtrip. If you find a document it does not handle correctly, file an issue with the offending file and I will look at it.

The deeper thesis

RTF is the format I started with because it is where my engineering history sits. The thesis is bigger than RTF.

The AI industry treats document ingestion as a preprocessing step rather than as the architectural foundation it actually is. The model is given the leftovers and asked to reconstruct what was thrown away. The answer is not a better model. The answer is preserving structure all the way through. Tables stay tables. Clauses stay clauses. Source pages stay referenceable. Diagnostics surface where confidence is low. The model reasons over the structured representation. Validation runs deterministically. The human review checkpoint sees what the system saw. The audit trail traces every step back to the source byte range in the original document.

This is what Sourcetrace is. rtfstruct handles the RTF case. pdfstruct handles PDFs. Other formats follow the same pattern. The commercial work I do at Lumen & Lever applies the same discipline at the architectural level: helping executives and boards establish control over AI before the structural mistakes compound.

The tools are free. Their purpose is adoption, technical credibility, and the slow accumulation of evidence that the structure-before-model thesis is right. Build with them. File issues. Fork them if you find a better path.

Try it

rtfstruct is on GitHub under Apache-2.0. Documentation is on GitHub Pages. The library installs from source today and will be on PyPI when the API stabilises.


Lee Powell is the architect of Scrivener and Scapple, a former enterprise architect at Commonwealth Bank and Deutsche Bank, and the founder of Lumen & Lever, an AI governance consultancy advising executives and boards on structural AI readiness.