惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

Microsoft Azure Blog
Microsoft Azure Blog
H
Hacker News: Front Page
A
About on SuperTechFans
云风的 BLOG
云风的 BLOG
aimingoo的专栏
aimingoo的专栏
Martin Fowler
Martin Fowler
博客园 - 叶小钗
Last Week in AI
Last Week in AI
Recent Announcements
Recent Announcements
P
Palo Alto Networks Blog
Webroot Blog
Webroot Blog
Hacker News: Ask HN
Hacker News: Ask HN
IT之家
IT之家
K
KPMG report finds enterprise disconnect between AI and its ROI | CIO
T
Threat Research - Cisco Blogs
C
CERT Recently Published Vulnerability Notes
Google DeepMind News
Google DeepMind News
Hugging Face - Blog
Hugging Face - Blog
H
Help Net Security
P
Privacy & Cybersecurity Law Blog
C
Cisco Blogs
罗磊的独立博客
The GitHub Blog
The GitHub Blog
M
MIT News - Artificial intelligence
人人都是产品经理
人人都是产品经理
The Cloudflare Blog
Y
Y Combinator Blog
AWS News Blog
AWS News Blog
H
Hackread – Cybersecurity News, Data Breaches, AI and More
钛媒体:引领未来商业与生活新知
钛媒体:引领未来商业与生活新知
K
Kaspersky official blog
博客园 - 司徒正美
Threat Intelligence Blog | Flashpoint
Threat Intelligence Blog | Flashpoint
Security Archives - TechRepublic
Security Archives - TechRepublic
The Last Watchdog
The Last Watchdog
Jina AI
Jina AI
MyScale Blog
MyScale Blog
TaoSecurity Blog
TaoSecurity Blog
大猫的无限游戏
大猫的无限游戏
cs.CL updates on arXiv.org
cs.CL updates on arXiv.org
Cisco Talos Blog
Cisco Talos Blog
美团技术团队
T
Tor Project blog
cs.CV updates on arXiv.org
cs.CV updates on arXiv.org
博客园 - 【当耐特】
博客园 - 聂微东
V2EX - 技术
V2EX - 技术
I
Intezer
V
Visual Studio Blog
酷 壳 – CoolShell
酷 壳 – CoolShell

DEV Community

Authentication Security Deep Dive: From Brute Force to Salted Hashing (With Java Examples) Why AI Systems Don’t Fail — They Drift Spilling beans for how i learn for exam😁"Reinforcement Learning Cheat Sheet" I Replaced Chrome with Safari for AI Browser Automation. Here's What Broke (and What Finally Worked) How Python Borrows Other People's Work The $40 Architecture: Processing 1 Billion API Requests with 99.99% Uptime Vibe Coding: A Workflow Guide (From Zero to SaaS) Most webhook security guides protect the wrong side. The scary part is delivery. Headless CMS for TanStack Start: Build a Blog with Cosmic EU Age Verification App "Hacked in 2 Minutes" — What Actually Happened Comfy Cloud’s delete function does not actually remove files Running AI Models on GPU Cloud Servers: A Beginner Guide Event-driven media intelligence with AWS Step Functions and Bedrock I scored 500 AI prompts across 8 quality dimensions — here's what broke How to Call Google Gemini API from Next.js (Free Tier, No Backend Needed) The Portal Protocol: Reclaiming Human Connection in the Age of AI How to Fix Your Team's Scattered Knowledge Problem With a Self-Hosted Forum Intro to tc Cloud Functors: A Graph-First Mental Model for the Modern Cloud Designing Multi-Tenant Backends With Both Ownership and Team Access I Built a Neumorphic CSS Library with 77+ Components — Here's What I Learned PostgreSQL Performance Optimization: Why Connection Pooling Is Critical at Scale Cómo construí un SaaS multi-rubro para gestionar expensas en Argentina con FastAPI + Vue 3 🚀 I Built an Ethical Hacking Scanner Tool – Open Source Project I Replaced /usage and /context in Claude Code With a Single Statusline A Pythonic Way to Handle Emails (IMAP/SMTP) with Auto-Discovery and AI-Ready Design I Collected 8.9 Million Polymarket Price Points — Here's What I Found About How Markets Really Move EcoTrack AI — Carbon Footprint Tracker & Dashboard Everyone's Using AI. No One Agrees How. 5 self-hosted ebook managers worth trying in 2026 Building Your First AI Agent with LangChain: From Chatbot to Autonomous Assistant Common SOC 2 Failures (Real World) Stop Vibe-Checking Your AI App: A Practical Guide to Evals How to Use SonarQube and SonarScanner Locally to Level Up Your Code Quality Your Next To-Do App Is Dead — I Replaced Mine with an OpenClaw AI Sign a Nostr event in 60 lines of Python using coincurve — no nostr-sdk, no nbxplorer, no rust toolchain ITGC Audit Explained Like You’re in Big 4 Patch Tuesday abril 2026: Microsoft parcha 163 vulnerabilidades y un zero-day en SharePoint Stop scraping everything: a better way to track competitor price changes Listing on MCPize + the Official MCP Registry while routing payments OUTSIDE the marketplace — how I kept 100% of my x402 revenue Building an AI-Powered Risk Intelligence System Using Serverless Architecture Why We Ripped Function Overloading Out of Our AI Toolchain Testing AI-Generated Code: How to Actually Know If It Works SaaS Churn Is Killing Your Business. Here Is What to Do About It (Without a Support Team) The Speed of AI Is No Longer Linear - And Self-Improving Models Are Why How to Implement RBAC for MCP Tools: A Practical Guide for Engineering Teams From Standard Quote to Persuasive Proposal: AI Automation for Arborists I built a CLI that scaffolds complete multi-tenant SaaS apps Axios CVE-2025–62718: The Silent SSRF Bug That Could Be Hiding in Your Node.js App Right Now The dashboard that ended our friendship Data Pipelines Explained Simply (and How to Build Them with Python) The Hidden Cost of AI Systems Nobody Talks About. undefined vs undeclared, and how typeof behaves Switching from file-based jobs to NATS/Kafka in Rust without changing code io_uring Adventures: Rust Servers That Love Syscalls Why Agentic AI is Killing the Traditional Database The POUR principles of web accessibility for developers and designers Quantum Neural Network 3D — A Deep Dive into Interactive WebGL Visualization How To Install Caveman In Codex On macOS And Windows Automation Pipeline Reliability: Why Your Workflow Breaks When Nobody Is Watching I Built an 'Open World' AI Coding Agent — It Works From ANY Folder From Freelancing to Product: A Tech Service Company's SaaS Transformation China's AI Giants: Adding Tencent Hunyuan & ByteDance Doubao to AI University (74 Providers) On the Vibe Coders and Their Lies clerk: Auto-Summarize Your Claude Code Sessions AI Weekly — 2026/04/10–04/17 | The Model Lockdown Is Here, but the Toolchain Is the Real Battleground AI 週報 — 2026/04/10–2026/04/17 模型封鎖潮來了,但工具鏈才是真戰場 Maybe this is how Open-Source apps are born... 🚀 Fine-Tune LLMs with LoRA and QLoRA: 2026 Guide tRPC v11 + Next.js App Router: End-to-End Type Safety Without the Boilerplate ShadCN UI in 2026: Why I Stopped Installing Component Libraries and Started Owning My Components SaaS Billing in React Server Components: Stripe + Supabase Without a Single `useEffect` Join our DEV Weekend Challenge — $1,000 in Prizes Across TEN winners! Submissions Due April 20 at 6:59 AM UTC. Implementing FSRS Spaced Repetition in Flutter + Supabase — Adding Memory Science to an AI Learning App "I Texted My Localhost From the Train — Claude Code Fixed the Bug Before I Got Home" I Built a Sales Prep AI and It Went Deeper Than Expected Design to Code #2: One JSON, Eleven Outputs Solving the 100M-Row Problem: A Summary Table Pattern for High-Volume Push Notification Logs Flutter Web With Wasm: What Actually Changes For Developers I Built 50 Royalty-Free Soundtracks for My Side Project in a Weekend Using AI Music Generation The Vibe Coding Security Checklist: 7 Things to Check Before You Ship Stop Letting Googlebot Guess Fix Your React App's SEO Right Desconstruindo o Streaming do LinkedIn: Como Criar um Engine de Extração de Vídeo de Alta Performance com HLS e FFmpeg (EDA Part-1) EDA (Exploratory Data Analysis) Explained With Real Life — Why Looking at Your Data Is the Most Important Step in Machine Learning Brand Relationship Management at Scale: Our 4-Touch Outreach System for 200+ Brands Why String.fromEnvironment() Might Return an Empty String in Dart JGuardrails 1.0.0 — Hardening Java LLM Apps Against Jailbreaks, Toxicity, and Prompt Injection Plan and Schedule a Full Week of Threads Content From One Claude Conversation Coding Cat Oran Ep3, Five Tables Changed Everything Updated: BFF Pattern I'm done watching freelancers get buried by 200 proposals. So I'm building the alternative. This is my first post BFS Algorithm in Java Step by Step Tutorial with Examples Tracking LLM Pricing Monthly: An Open Dataset for 22 AI Models How We Measure Content ROI on a Comparison Site: Revenue Attribution Without Perfect Data Introducing Nova AI Ops: The AI-Native Operating System for SRE Teams I built a free desktop video downloader for Windows — Grabbit How Talkie OCR Helps Vision-Impaired & Dyslexic Users Read the World Around Them VRCFaceTracking安装和iPhone面捕配置教程,有bug Even CrowdStrike Can't See Your Agents The Automation Gold Rush: What n8n Workflows and Claude Are Opening Up for Developers Right Now
AI Document Processing in Production: Full Pipeline Guide
Iurii Rogulia · 2026-06-24 · via DEV Community

Someone emails you a PDF invoice. You want to extract the vendor name, line items, total amount, currency, and due date — automatically, at scale, without manual keying.

You call the OpenAI API, pass the PDF as base64, get a JSON blob back. It works. You ship it. Then reality arrives: a scanned invoice from a vendor who still uses a physical stamp. A 60-page contract where the key clause is on page 47. A table-heavy bank statement where amounts bleed across column boundaries. A PDF that's actually an image with no embedded text at all.

The naive approach collapses on all of them. Here's the production architecture that does.

Why the Naive Approach Breaks

The simplest version — encode the whole PDF, send it to GPT, ask it to return JSON — fails in four common ways:

Token limits. A 50-page contract is roughly 25,000–40,000 tokens of text, plus image tokens if you're sending page renders. Most model context windows handle it technically, but accuracy degrades on long documents. The model loses track of structure. Extraction quality on page 45 is noticeably worse than page 2.

Scanned documents. A PDF with no embedded text layer is just a sequence of images. No amount of prompting extracts text that isn't there. You need OCR. This affects more documents than you expect — expense receipts, legacy contracts, anything printed and scanned, anything generated by certain accounting systems.

Tables. Tables are the hardest part of PDF extraction. Embedded text in a PDF doesn't encode column relationships — the text objects are just positioned by x/y coordinates. A naive extraction reads the text linearly and loses the table structure entirely. Line items from an invoice become a flat list with no mapping between description, quantity, and amount.

Cost at scale. Sending a 10MB PDF as a single API call costs real money and consumes tokens inefficiently. Most of that context is headers, footers, boilerplate legal text, and page numbers. The fields you actually need are in 5% of the document.

The Full Pipeline

Input PDF
  → File validation (size, type, not encrypted)
  → Text extraction attempt (pdfplumber / pypdf)
  → Quality check: does extracted text look usable?
  → If not: OCR pipeline (Tesseract / Textract / Document AI)
  → Page splitting + relevant-page detection
  → Table extraction (if document type warrants it)
  → Structured model call (JSON mode / tool use)
  → Output validation (schema check + cross-field rules)
  → Confidence scoring
  → Storage or human-review queue

Each stage has a failure mode. Each one needs its own handling.

slug="ai-integration"
text="Building AI document processing for EU invoices, contracts, or compliance workflows? I design and ship these pipelines end-to-end."
/>

Stage 1: Text Extraction

Start with embedded text — it's faster, cheaper, and more accurate than OCR when available.

I use pdfplumber for extraction in Python. It handles text positioning better than pypdf and gives you bounding box data that's useful for table detection:

# extract.py
from __future__ import annotations

import pdfplumber

def extract_text_by_page(pdf_path: str) -> list[dict]:
    """
    Extract text from each page. Returns a list of dicts with
    page number, raw text, and a usability flag.
    """
    pages = []

    with pdfplumber.open(pdf_path) as pdf:
        for i, page in enumerate(pdf.pages):
            text = page.extract_text(x_tolerance=3, y_tolerance=3) or ""
            word_count = len(text.split())

            pages.append({
                "page": i + 1,
                "text": text,
                "word_count": word_count,
                # Heuristic: fewer than 30 words on a non-blank page = likely scanned
                "needs_ocr": word_count < 30 and len(page.images) > 0,
            })

    return pages

def extract_tables(pdf_path: str, page_number: int) -> list[list[list[str]]]:
    """
    Extract tables from a specific page. Returns a list of tables,
    each table being a list of rows, each row being a list of cell strings.
    """
    with pdfplumber.open(pdf_path) as pdf:
        page = pdf.pages[page_number - 1]
        tables = page.extract_tables()
        # Normalize: replace None cells with empty string
        return [
            [[cell or "" for cell in row] for row in table]
            for table in (tables or [])
        ]

The per-page needs_ocr flag is the key decision point. A hybrid document — mostly text with one scanned attachment page — gets OCR applied only to the scanned pages, not the whole file.

Stage 2: OCR When You Need It

Three options, each with different tradeoffs:

Tesseract — open source, free, runs locally. Accuracy is acceptable for clean scans, poor for low-resolution, skewed, or multi-column layouts. Good default for low-volume pipelines where you control the document source.

AWS Textract — purpose-built for documents. Handles tables natively, returns structured output with bounding boxes and confidence scores per word. Costs $0.0015 per page for basic detection, $0.015 for table/form extraction. The table extraction is worth the cost for invoices and financial statements.

Google Document AI — strongest accuracy for complex layouts and multilingual documents. More expensive than Textract, but noticeably better on documents with mixed scripts or unusual formatting.

My heuristic: Tesseract for internal tooling and prototypes, Textract for invoices and financial documents, Document AI if you're processing government documents or multi-language content.

For Textract integration from Python:

# ocr_textract.py
from __future__ import annotations

import boto3

def ocr_page_with_textract(image_bytes: bytes) -> dict:
    """
    Run a single page image through AWS Textract.
    Returns raw blocks with type, text, and confidence.
    """
    client = boto3.client("textract", region_name="eu-west-1")
    response = client.detect_document_text(
        Document={"Bytes": image_bytes}
    )

    lines: list[dict] = []
    for block in response["Blocks"]:
        if block["BlockType"] == "LINE":
            lines.append({
                "text": block["Text"],
                "confidence": block["Confidence"],
                "bbox": block["Geometry"]["BoundingBox"],
            })

    return {
        "lines": lines,
        "full_text": "\n".join(b["text"] for b in lines),
        "min_confidence": min((b["confidence"] for b in lines), default=0),
    }

For table extraction specifically, use analyze_document with FeatureTypes=["TABLES"] — this is a separate call and a higher per-page cost, but it's the only reliable way to get structured table data from scanned documents.

Stage 3: Relevant Page Detection

For a 60-page contract, you don't send all 60 pages to the model. You find the pages that contain the data you need.

The approach: keyword scoring per page. For invoice extraction, I score pages by the presence of terms like "invoice", "total", "amount", "due date", "vendor", and their localized equivalents. The top-scoring 3–5 pages get sent to the model. For contracts, I look for "payment terms", "effective date", "party", "agrees to".

# relevance.py
from __future__ import annotations

import re

INVOICE_KEYWORDS = [
    "invoice", "total", "amount due", "subtotal", "tax", "vat",
    "due date", "payment terms", "bill to", "vendor", "supplier",
    # Finnish (common in EU processing)
    "lasku", "summa", "eräpäivä", "alv",
]

def score_page_relevance(text: str, keywords: list[str]) -> float:
    """
    Returns a 0.0–1.0 relevance score for a page given target keywords.
    """
    if not text:
        return 0.0

    text_lower = text.lower()
    matches = sum(1 for kw in keywords if kw in text_lower)
    return min(matches / max(len(keywords) * 0.3, 1), 1.0)

def select_relevant_pages(
    pages: list[dict],
    keywords: list[str],
    max_pages: int = 5,
) -> list[dict]:
    scored = [
        {**page, "relevance": score_page_relevance(page["text"], keywords)}
        for page in pages
    ]
    scored.sort(key=lambda p: p["relevance"], reverse=True)
    return [p for p in scored[:max_pages] if p["relevance"] > 0.05]

This alone cuts token consumption by 70–80% on long documents while keeping extraction quality the same or better — because the model isn't distracted by irrelevant content.

Keyword scoring is a pragmatic baseline — for higher recall and precision, embedding-based retrieval per page can replace it, but at higher cost and complexity.

Stage 4: Structured Extraction

This is where most tutorials go wrong. They send a prompt like "extract the invoice fields and return JSON". The model returns JSON most of the time. Sometimes it returns JSON wrapped in a markdown code fence. Sometimes it adds commentary before the JSON. Sometimes it invents fields. You end up writing a fragile parser on top of an unpredictable output.

Use JSON mode (OpenAI) or tool use. These are not optional conveniences — they're the difference between a system that works reliably and one that works most of the time.

Here is the TypeScript extraction call using the Vercel AI SDK with a Zod schema enforced via tool use:

// lib/documents/extract-invoice.ts
import { openai } from "@ai-sdk/openai";
import { generateObject } from "ai";
import { z } from "zod";

const InvoiceSchema = z.object({
  vendor_name: z.string().describe("The name of the company issuing the invoice"),
  vendor_vat_number: z
    .string()
    .nullable()
    .describe("VAT registration number if present, null otherwise"),
  invoice_number: z.string().describe("The invoice reference number"),
  invoice_date: z.string().describe("Date the invoice was issued, ISO 8601 format if possible"),
  due_date: z
    .string()
    .nullable()
    .describe("Payment due date, ISO 8601 format if possible, null if not found"),
  currency: z.string().describe("ISO 4217 currency code, e.g. EUR, USD, GBP"),
  subtotal: z.number().nullable().describe("Pre-tax amount as a number, null if not found"),
  tax_amount: z.number().nullable().describe("Tax/VAT amount as a number, null if not found"),
  total_amount: z.number().describe("Total amount due as a number"),
  line_items: z.array(
    z.object({
      description: z.string(),
      quantity: z.number().nullable(),
      unit_price: z.number().nullable(),
      total: z.number().nullable(),
    })
  ),
  confidence: z.number().min(0).max(1).describe("Your confidence in this extraction, 0.0 to 1.0"),
});

export type InvoiceExtraction = z.infer<typeof InvoiceSchema>;

export async function extractInvoice(
  pageTexts: string[],
  tableData: string[][][][] = []
): Promise<InvoiceExtraction> {
  const context = pageTexts.join("\n\n---\n\n");
  const tableContext =
    tableData.length > 0
      ? "\n\nExtracted tables:\n" +
        tableData.map((t) => t.map((r) => r.join(" | ")).join("\n")).join("\n\n")
      : "";

  const { object } = await generateObject({
    model: openai("gpt-4o-mini"),
    schema: InvoiceSchema,
    prompt: `Extract all invoice fields from the following document text.
Use null for fields not present in the document.
For currency, always return an ISO 4217 code.
For dates, normalize to YYYY-MM-DD if possible.
Set confidence to reflect how certain you are about the extraction overall.

Document:
${context}${tableContext}`,
  });

  return object;
}

The confidence field in the schema is deliberate — I ask the model to self-report its certainty. It's not always calibrated perfectly, but it's a useful first filter. Extractions with confidence < 0.6 go to the human review queue automatically.

Stage 5: Validation

Structured output from the model is not validated output. The model will comply with the schema — but it can still produce values that pass schema validation while being logically wrong.

// lib/documents/validate-invoice.ts
import type { InvoiceExtraction } from "./extract-invoice";

export interface ValidationResult {
  valid: boolean;
  errors: string[];
  warnings: string[];
}

export function validateInvoiceExtraction(data: InvoiceExtraction): ValidationResult {
  const errors: string[] = [];
  const warnings: string[] = [];

  // Hard failures — this extraction is unreliable
  if (!data.vendor_name || data.vendor_name.trim().length < 2) {
    errors.push("vendor_name is missing or too short");
  }

  if (data.total_amount <= 0) {
    errors.push("total_amount must be greater than zero");
  }

  if (!data.currency || !/^[A-Z]{3}$/.test(data.currency)) {
    errors.push(`currency '${data.currency}' is not a valid ISO 4217 code`);
  }

  // Cross-field logic
  if (data.subtotal !== null && data.tax_amount !== null) {
    const expectedTotal = data.subtotal + data.tax_amount;
    const delta = Math.abs(expectedTotal - data.total_amount);
    if (delta > 0.02) {
      // Tolerance for rounding
      errors.push(
        `subtotal (${data.subtotal}) + tax (${data.tax_amount}) = ${expectedTotal}, ` +
          `but total_amount is ${data.total_amount}. Delta: ${delta.toFixed(2)}`
      );
    }
  }

  if (data.total_amount > 0 && !data.currency) {
    errors.push("total_amount is present but currency is missing");
  }

  // Soft warnings — flag for review but don't reject
  if (data.line_items.length === 0) {
    warnings.push("no line items extracted — verify manually");
  }

  if (data.confidence < 0.6) {
    warnings.push(`low model confidence: ${data.confidence}`);
  }

  if (!data.due_date) {
    warnings.push("due_date not found — may need manual entry");
  }

  const vatPattern = /^[A-Z]{2}\d{8,12}$/;
  if (data.vendor_vat_number && !vatPattern.test(data.vendor_vat_number.replace(/\s/g, ""))) {
    warnings.push(`vendor_vat_number '${data.vendor_vat_number}' doesn't match expected format`);
  }

  return {
    valid: errors.length === 0,
    errors,
    warnings,
  };
}

The subtotal + tax = total cross-check catches more extraction errors than any other single rule. Models sometimes extract the subtotal as the total, or miss the tax component entirely.

Stage 6: Low-Confidence Handling

Two strategies, not mutually exclusive:

Automated retry with a stronger model. If confidence < 0.7 and validation passes but has warnings, retry with gpt-4o instead of gpt-4o-mini. The cost difference is roughly 15×, so don't do this for all documents — only for those where the faster model flagged uncertainty. In practice this applies to 10–15% of documents and catches most of the edge cases.

Human-in-the-loop queue. If validation returns errors, or if the retry also produces low confidence, route to a review interface. The key is to pre-fill the UI with the extracted values — the reviewer confirms or corrects, they don't start from scratch. This makes human review fast enough to be operationally viable even at moderate volume.

// lib/documents/pipeline.ts
import { extractInvoice } from "./extract-invoice";
import { validateInvoiceExtraction } from "./validate-invoice";
import { db } from "@/lib/db";

type ProcessingOutcome = "stored" | "retry" | "human_review";

export async function processInvoiceDocument(
  pageTexts: string[],
  tables: string[][][][],
  documentId: string
): Promise<ProcessingOutcome> {
  let result = await extractInvoice(pageTexts, tables);
  let validation = validateInvoiceExtraction(result);

  // Retry with stronger model if first pass is uncertain
  if (!validation.valid || result.confidence < 0.7) {
    result = await extractInvoice(pageTexts, tables); // gpt-4o retry handled inside
    validation = validateInvoiceExtraction(result);
  }

  if (!validation.valid) {
    await db.insert(reviewQueue).values({
      documentId,
      extractedData: result,
      validationErrors: validation.errors,
      validationWarnings: validation.warnings,
      status: "pending_review",
    });
    return "human_review";
  }

  await db.insert(invoices).values({
    documentId,
    ...result,
    processedAt: new Date(),
  });
  return "stored";
}

Throughput and Cost at Scale

At low volume (under 100 documents/day), synchronous processing is fine. At scale, you need async. If this is part of a broader automation workflow, BullMQ fits naturally as the processing backbone.

Queue documents with BullMQ. Set concurrency based on your rate limits — OpenAI's Tier 2 allows 5,000 RPM for gpt-4o-mini. With 3 API calls per document (extraction + possible retry + any enrichment), you can process roughly 1,600 documents per minute at full throughput before hitting rate limits.

Model selection by document type matters for cost:

Document type Model Avg cost/doc Notes
Clean digital invoice gpt-4o-mini ~$0.004 High accuracy, fast
Complex table-heavy doc gpt-4o ~$0.06 On retry/escalation only
Scanned + OCR needed Textract + 4o-mini ~$0.018 Textract adds $0.015/page
Contract (10+ pages) gpt-4o-mini ~$0.012 After page filtering

The page relevance filtering from Stage 3 is the biggest cost lever. A 20-page document with 3 relevant pages costs 85% less than sending all 20 pages.

Monitoring

What breaks silently: extraction quality drift. The model doesn't throw an error — it just starts extracting total_amount less reliably as your document corpus evolves and new formats appear.

Log these metrics per processing run:

  • Validation error rate (errors / total documents)
  • Human review queue depth and queue age
  • Model confidence score distribution (P25, P50, P75)
  • Retry rate (documents requiring the stronger model)
  • Cross-field validation failure breakdown by rule

Set an alert if the validation error rate crosses 5% in a rolling 1-hour window. That's your signal that a new document format is breaking the pipeline and needs a schema or prompt update.

Gotchas From Production

pdfplumber hangs on encrypted PDFs. If the document is password-protected, the extraction call never returns. Add a timeout and catch the exception. Check for encryption before attempting extraction.

Scanned PDFs often have a text layer that's garbage. Some scan-to-PDF workflows produce PDFs with embedded text, but the text is from a failed OCR pass — garbled characters, incorrect word boundaries. My heuristic for detecting this: extract text, then count the ratio of alphabetic characters to total characters. Below 0.6, treat it as scanned regardless of what pdfplumber returns.

The model sometimes returns numbers with thousand separators. "1,234.56" — and Zod will reject it because the schema expects a number. Normalize numeric strings before schema validation: strip commas, handle both . and , as decimal separators (European invoices use comma as decimal).

Date formats vary wildly. 14.03.2025, 14/03/2025, March 14, 2025, 2025-03-14. The model normalizes most of them, but it's worth adding a post-processing step that parses the extracted date string through date-fns/parse with multiple format attempts before storing.

Textract is slow for real-time use. Average latency is 3–8 seconds per page for detect_document_text. If you need synchronous responses, use Tesseract locally for a fast first pass and Textract in an async enrichment step.

Results

I built this pipeline for internal document processing and as a reusable module across client projects. In production use:

  • 94% of clean digital invoices extracted automatically with no human review
  • 78% of scanned invoices extracted automatically after the OCR stage
  • Validation cross-checks catch ~60% of extraction errors before they reach storage
  • Average cost per document: $0.005–$0.018 depending on document complexity
  • Human review queue processes the remaining 6–22% depending on document source quality

If you're building document processing automation for EU businesses — invoice processing, contract extraction, compliance document workflows — the pipeline above is where you start. The naive "pass to GPT" approach works in a demo. Production requires the full stack: pre-processing, OCR fallbacks, structured schemas, validation, and a human-in-the-loop safety net for the edge cases.

I've deployed this end-to-end across several production systems, including htpbe.tech (PDF forensics) and document automation pipelines for e-commerce and logistics clients.

Document processing is not an AI problem — it's a systems problem with an AI component.

If you need a senior developer who can build AI document processing that actually works at scale — get in touch. I'm available for freelance projects and long-term engagements.


Related reading: How I Detect Tampered PDFs in 9 Seconds — forensic analysis of PDF structure for document authenticity verification. How to Add AI to an Existing Product Without Rewriting It — the three integration patterns and where document processing fits.