惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

Security Latest
Security Latest
Know Your Adversary
Know Your Adversary
S
Schneier on Security
K
Kaspersky official blog
cs.CL updates on arXiv.org
cs.CL updates on arXiv.org
Google DeepMind News
Google DeepMind News
Project Zero
Project Zero
Recorded Future
Recorded Future
T
The Blog of Author Tim Ferriss
P
Palo Alto Networks Blog
Engineering at Meta
Engineering at Meta
MyScale Blog
MyScale Blog
L
LINUX DO - 热门话题
The Register - Security
The Register - Security
B
Blog
V
Vulnerabilities – Threatpost
M
MIT News - Artificial intelligence
P
Proofpoint News Feed
Cyberwarzone
Cyberwarzone
Latest news
Latest news
Webroot Blog
Webroot Blog
Cyber Security Advisories - MS-ISAC
Cyber Security Advisories - MS-ISAC
美团技术团队
奇客Solidot–传递最新科技情报
奇客Solidot–传递最新科技情报
C
CXSECURITY Database RSS Feed - CXSecurity.com
L
LINUX DO - 最新话题
K
KPMG report finds enterprise disconnect between AI and its ROI | CIO
B
Blog RSS Feed
aimingoo的专栏
aimingoo的专栏
The GitHub Blog
The GitHub Blog
D
Docker
G
Google Developers Blog
T
Threatpost
F
Full Disclosure
Attack and Defense Labs
Attack and Defense Labs
L
Lohrmann on Cybersecurity
Application and Cybersecurity Blog
Application and Cybersecurity Blog
月光博客
月光博客
D
DataBreaches.Net
WordPress大学
WordPress大学
F
Fortinet All Blogs
S
Secure Thoughts
Y
Y Combinator Blog
博客园 - Franky
Scott Helme
Scott Helme
Apple Machine Learning Research
Apple Machine Learning Research
Hacker News: Ask HN
Hacker News: Ask HN
酷 壳 – CoolShell
酷 壳 – CoolShell
Cloudbric
Cloudbric
N
News | PayPal Newsroom

Inside Nutrient

A guide to the invisible work behind documents Introducing Nutrient Documents for Salesforce: Native document generation and signing Document AI vs. traditional OCR: Choosing between OCR, AI, and hybrid pipelines PDF SDK compliance and security evaluation checklist for enterprise teams (2026) Invariant Corp replaces paper processes with Nutrient Workflow and scales without limits What is process mapping? A complete guide Nutrient vs. Conga Composer for Salesforce document generation (2026) Document routing: How to automate document distribution The CTO’s AI playbook: Why accountability architecture beats orchestration Compliance workflow automation: Why built-in compliance is table stakes Workflow diagrams: Examples, symbols, and how to build one that actually runs Digital forms: Replace paper forms with automated workflows Approval workflow software: How to automate approvals Why document-centric automation is different The CEO’s AI playbook: Why decision architecture beats model selection Nutrient SDK product updates for Q1 2026 PDF redaction verification: How to prove sensitive data is permanently removed What is a VPAT? The complete guide to accessibility conformance reports What is PDF/UA? The accessible PDF standard explained Salesforce eSignatures: Generate, sign, and track documents in one flow Online document viewer: Options, tradeoffs, and how to embed one Document viewer for web apps: React, Vue, Angular (2026) Best document viewers in 2026: A buyer’s guide How to edit a PDF in Python: Add text, images, and annotations Nutrient advances Workflow platform with agentic AI for enterprise-grade speed and consistency in document-heavy operations How to create a Salesforce quote template from opportunity data The business case for accessibility: Five ways it drives enterprise value Python PDF library comparison (2026): 7 libraries for developers Why your AI agent hallucinates PDF table data PDF.js limitations: When to upgrade to a commercial PDF SDK How Subject scaled 5× with Nutrient’s PDF SDK without rebuilding its document layer I replaced our sales training with an AI coach that runs in Slack — here’s what broke Redirecting to: https://securitybuzz.com/cybersecurity-news/why-enterprise-permissions-are-ais-most-dangerous-inheritance/ Nutrient .NET SDK vs. iText Core: Complete comparison for .NET developers DocuVieware: Support’s most frequently asked setup questions Introducing Nutrient Workflow How to convert PDF to Word in C# (.NET) When email and spreadsheets stop working: Work order approval workflows for field teams on the move Compliance with confidence: Why document-centric automation is the foundation of your mission Nutrient expands AI Assistant, automating multistep document workflows inside any application What is document generation? A developer’s guide to PDF generation Document Converter data flow and how real-time watermarks skip the queue PDF/UA compliance guide: Requirements, standards, and best practices How Athena Intelligence built AI agents for regulated enterprises with Nutrient’s document infrastructure How to convert HTML to PDF (2026): 4 methods from browser print to SDK How to build a document extraction pipeline with Nutrient Vision API OCR vs. intelligent document processing: Choosing the right document extraction engine Beyond OCR: How document intelligence eliminates manual processing in regulated industries Nutrient vs. IronPDF: Complete comparison for .NET developers Nutrient vs. Aspose.PDF: Complete comparison for .NET developers Redirecting to: https://fortune.com/2026/02/19/openclaw-who-is-peter-steinberger-openai-sam-altman-anthropic-moltbook/ Lufthansa Systems uses Nutrient to deliver reliable, scalable PDF rendering for pilots worldwide Nutrient vs. Syncfusion: Complete comparison for .NET developers React’s useTransition: The hook you’re probably using wrong First City Monument Bank streamlines banking processes with Nutrient Workflow Redirecting to: https://www.sdcexec.com/warehousing/automation/article/22957364/nutrient-workflow-automation-the-missing-link-in-supply-chain-efficiency The complete guide to digital signatures: PAdES, CAdES, and XAdES explained Nutrient Python SDK: Production-grade document processing for Python Introducing agentic document editing for web applications with AI Assistant Nutrient vs. QuestPDF: Complete comparison for .NET developers How we fixed the GdPicture license expiration (and what to do if you’re affected) Red team security testing with agentic AI The future of healthcare document automation Best healthcare workflow software compared Nutrient SDK product updates for Q4 2025 How Harvey scaled legal document workflows 50 percent MoM without rebuilding infrastructure HIPAA-compliant document management in hospitals How we optimized rendering performance while handling thousands of annotations in React — Part 2 Automated PII removal with Nutrient API Redirecting to: https://www.devopsdigest.com/2026-low-code-no-code-predictions Redirecting to: https://www.kmworld.com/Articles/Editorial/ViewPoints/Leaders-predict-AI-to-continue-permeating-all-aspects-of-KM-in-2026-172594.aspx What are deep agents and how do they solve complex problems? Whipping up document magic: Your easy-bake recipe for Vue and Nutrient Web SDK 🧁 What I’ve learned about product iteration planning while building SDKs Passwordless document signing: Three-layer security guide New zip folder functionality streamlines file management in Document Automation Server The keyboard shortcuts playbook: Taking control of keyboard events in Nutrient Web SDK From experienced engineer to AI beginner: My unexpected journey AI-assisted manual testing: Handling Safari’s PDF rendering and UI quirks How to keep a 20-year-old SDK up to date How we optimized rendering performance while handling thousands of annotations in React — Part 1 Nutrient announces new executive hires to accelerate next phase of growth High performance UI using web workers Automate document conversion at scale with Python and Nutrient DCS From curiosity to PLG (and AI): My journey to understanding product-led growth Prost to progress: One year as Nutrient Pigeon usage at Nutrient: Bridging native SDKs to Flutter Modernizing CI build servers: How to migrate from Chef to Ansible Unix man pages: AI-friendly documentation since 1971 Consistent hashing for even load distribution Best AI redaction APIs: Complete comparison guide for 2025 Why AI document redaction matters for modern security From coding to coordinating: How AI transformed my workflow What is intelligent document processing (IDP)? A complete guide Enterprise PDF SDKs: Best PSPDFKit (now Nutrient) alternatives Nutrient SDK product updates for Q3 2025 GdPicture support best practices Redacting sensitive data with Nutrient AI redaction API How AI is transforming the customer experience at Nutrient: From instant answers to intelligent support How manual QA uses PR testing between releases
Computers still can’t understand you
Chris Van Wesep · 2026-03-13 · via Inside Nutrient

Every so often, two pieces of news collide in your feed in a way that makes you stop scrolling. This week, I got one of those moments.

Two articles landed in my feed back-to-back, and they made me wonder if the universe isn’t actually some kind of cosmic simulation after all.

The first was from The Tech Buzz, and it was entitled “AI’s Dirty Secret: It Still Can’t Read PDFs Properly(opens in a new tab).” Not an unusual topic for that particular source. The second, though, really forced a double-take. It was published in The Economist and it was titled “The war against PDFs is heating up(opens in a new tab).”

The Economist — that same publication that talks about global trade, interest rates, international conflict, and neoliberal macroeconomic theory, talking about a software document format? What in Turing’s good name is going on?

For those who haven’t been paying close attention to US domestic news, the federal Department of Justice just released more than three million pages of PDF documents for public review. Like many government documents, many of these weren’t originally software-generated; a great many were still hand-typed — or worse, handwritten — documents that needed translation into machine-readable format.

What emerged comes as no real surprise to anyone who’s ever had to deal with this before: It was a mess. The basic optical character recognition (OCR) technology used for processing those source documents produced garbled, unsearchable output. Not only did the public have reason to wonder at the technical competence of the DOJ staffers, but journalists and researchers couldn’t find what they were looking for in documents the government was legally required to make transparent.

And yet, here’s the thing that caught my attention: This wasn’t treated as a technical glitch. It was treated as a revelation — as if the industry had just discovered that extracting structured data from PDFs is still an unsolved problem.

This problem isn’t new. The industry has been quietly struggling with it since the ‘90s. And the team here at Nutrient has been quietly solving it for more than a decade.

The Economist piece posed the question: Can developers build tools that handle PDFs properly? TL;DR: Yes, but like most formats, you have to know what you’re working with when you’re working with a PDF.

Why this is structurally difficult

To understand why document extraction is so hard, you need to understand what a PDF actually is — or, perhaps, what it isn’t.

Consider HTML for a moment. An HTML document contains within it an intentionally explicit structure to help renderers. This is why the HTML specification describes “semantic” HTML markup, as the tag language is designed to reflect the semantics rather than just display characteristics. A bulleted list is bookended by a starting and ending <ul> tag. A table is set off by a <table> tag, and each row and cell is set off by additional tag pairs. Even that most humble of structural elements — the <div> — tells the HTML renderer that something is a new section of markup.

On the other hand, PDF, like its immediate predecessor, PostScript, is a set of printing instructions, a la code, that tells a renderer where to place ink on a page or screen. It doesn’t encode what a table is, where a paragraph begins, or what reading order the content follows. A PDF generated from a Word document often captures none of the underlying structural elements like headers, footers, or page numbers. The structure is obvious in the human’s eyes, but to the PDF, it’s all just formatting commands that happen to look correct when rendered visually. Scanned PDFs are worse: They’re literally arrays of bytes that happen to form images, with no machine-readable text at all.

This is why the DOJ’s documents came out garbled. The OCR step converted pixels to characters just fine — the words were all there. But the extraction pipeline treated the result as a flat text stream, and every bit of structural context was lost.

This is a well-understood failure pattern, and it shows up in predictable ways:

Table extraction failure. Tables in PDFs are often just aligned text and drawn lines. Without structure-aware parsing, columns collapse, headers detach from data, and multi-row cells merge incorrectly.

Multicolumn layout failure. Government documents, academic papers, and annual reports use multicolumn layouts. Naive extraction reads left-to-right across both columns simultaneously, producing nonsense.

Downstream hallucination. When extraction produces partial or disordered text and that text is fed to an LLM, the model fills in the gaps. On financial documents, contracts, or medical records, that means plausible-but-wrong numbers and dates.

The good news is that the industry has been making progress — both with specialized document extraction tools and with multimodal AI models that can process pages as images rather than raw text. The bad news is that each approach comes with its own tradeoffs in quality, infrastructure control, and cost.

Why “better OCR” isn’t the answer — and why multimodal AI isn’t either

The industry’s first instinct is to throw more OCR at the problem. More accuracy, more languages, more preprocessing. But as the DOJ case showed, character recognition accuracy wasn’t the issue — the words were all there. The problem is structural understanding: knowing that these cells form a table, that this paragraph precedes that one, that this handwritten annotation is a dosage linked to a patient field.

The second instinct is to skip the extraction pipeline entirely and point a multimodal AI model at the page. Claude, Gemini, and ChatGPT can all process PDF pages as images now, and they’re meaningfully better than naive text extraction. But they have their own limitations: They struggle with merged table cells, truncate multipage tables at page breaks, produce inconsistent results across runs, and provide no confidence scores to tell you what’s reliable and what isn’t. Real-world tests have found up to 42 percent of fields missing from LLM-extracted data on complex documents. And they return Markdown or text, not structured data with spatial coordinates you can trace back to the source page.

Here’s a useful framing. Think of document extraction as three tiers.

Tier 1: Character recognition. Converting pixels to text. Most tools do this reasonably well. It’s fast and lightweight, but it’s all you get — a flat string of characters with no structural context.

Tier 2: Structural understanding. Determining reading order, detecting columns, extracting tables with cell-level coordinates, recognizing form fields, preserving hierarchical relationships between document elements. This is where most tools stop — and where most failures originate. It’s also what makes extracted data actually usable downstream.

Tier 3: AI-enhanced analysis. Running OCR, intelligent content recognition (ICR), and a vision language model in parallel, and then merging the results — combining the spatial precision of specialized local models with the general document understanding of VLMs for the hardest documents: irregular table layouts, degraded scans, and complex handwriting.

So what’s left? Cloud-only platforms like Adobe’s Acrobat AI assistant and Google Document AI are pushing into tiers 2 and 3 with dedicated document extraction APIs. These are more reliable than raw LLM processing — but they require every document to leave your infrastructure. For regulated industries — healthcare, legal, financial services, government — that’s often a non-starter before the conversation even begins. They also come with per-page pricing that scales linearly with volume, and as the recent software outages at OpenAI and Microsoft have shown, dependency on external services brings a degree of fragility that many organizations cannot permit.

The question isn’t whether AI can help with document extraction — it can. The question is where that AI runs, who controls it, and whether your extraction pipeline works without it.

What structural document intelligence looks like

One of the reasons I joined Nutrient as CMO is that the people here know PDF really, really well. The engineering team has been building PDF processing and document understanding technology for more than a decade. That expertise now powers Vision API, which launched this week as part of our Python and Java SDKs.

Vision API isn’t an incremental improvement to OCR; it’s a different architecture. Three modes of operation — OCR, ICR, and VLM-enhanced ICR — are each designed for a different level of the extraction problem, and they’re all accessible through a single API. You pick the operation mode that matches the document.

The OCR mode handles tier 1: fast character recognition with word-level bounding boxes.

ICR is where it gets interesting. It uses specialized on-premises AI models — small models that are each optimized for a specific task, like segmentation, character recognition, or table detection — to handle tier 2: document layout detection, table extraction with cell-level coordinates, equation recognition, handwriting, and correct reading order. It’s all local, and nothing leaves your infrastructure. For most documents — standard tables, forms, two-column layouts, mixed print, and handwriting — ICR handles the job on its own. No cloud calls, no external dependencies.

For the most difficult documents — degraded scans, complex handwriting, unusual layouts — developers can opt into VLM-enhanced ICR (tier 3). Here’s how it works: All three engines — OCR, ICR, and a vision language model (Claude, OpenAI, or a locally hosted model) — run in parallel on the presegmented document. OCR contributes exact character recognition. ICR contributes spatial precision, bounding boxes, and structural layout. The VLM contributes general document understanding that specialized models lack — it’s better at handling unusual layouts where text appears in unexpected positions. The results are then merged into a single cohesive output that’s more accurate than any engine alone. You choose whether and when data leaves your environment — for full isolation, you can use a local VLM like Qwen instead of a cloud provider.

This is what separates Vision API from the general LLM approach described earlier. Where LLMs give you Markdown or text with no spatial context, Vision API returns structured JSON for every element — with its type, bounding box coordinates, and position in the reading order. You can trace any value back to the exact pixel region in the original page. That’s what makes audit trails, review UIs, and compliance reporting possible. And because the merged output combines ICR’s spatial precision and bounding boxes with VLM’s general document understanding, you avoid the positioning errors, hallucinated text, and missing fields that plague VLM-only extraction.

It’s important that developers know where and how the document is being processed, but it’s business-critical to the CEO and the CISO. For the healthcare company handling patient records, the bank processing loan documents, the law firm reviewing discovery materials, and the government agency managing classified files, sending every document to a cloud API isn’t viable. In fact, it’s often disqualifying at procurement.

The right approach isn’t to choose between local processing and cloud AI. It’s to run the right AI at the right layer. Vision API handles extraction — the structural engineering problem — using on-device AI that never leaves your infrastructure. Combining ICR’s spatial precision with VLM’s general document understanding produces better results than either engine alone — and the structured JSON output feeds downstream LLMs far better input than a flat text dump or Markdown. Better input produces better output, at every layer.

Get started with Vision API using our Python or Java guides, or contact our team to discuss your document processing pipeline.

The bigger picture

The Tech Buzz called document processing “the stuff that actually matters for day-to-day business operations.” It’s infrastructure. It’s not glamorous. It doesn’t generate headlines about artificial general intelligence. But an estimated 2.5 trillion PDFs exist in the world, and every industry depends on them. (And, as any reader of XKCD knows, introducing a new standard to replace all the other ones just adds a new one to the pile.)

The organizations that solve document intelligence at the extraction layer — with control over where it runs and what it costs — unlock everything downstream: search, compliance, automation, AI-powered analysis.

One startup profiled by The Economist is trying to build a new file type to replace the PDF entirely. That’s a bold bet against 30 years of institutional adoption. PDFs aren’t going anywhere, and building a new format isn’t the challenge before us. It’s to build tools that understand them structurally, run locally, and give developers control over the accuracy-privacy tradeoff.

That’s what the team here has been doing for more than a decade — and it’s the story I want more people to hear.