惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

Cyberwarzone
Cyberwarzone
Jina AI
Jina AI
WordPress大学
WordPress大学
N
Netflix TechBlog - Medium
freeCodeCamp Programming Tutorials: Python, JavaScript, Git & More
Google DeepMind News
Google DeepMind News
博客园 - 司徒正美
宝玉的分享
宝玉的分享
C
Check Point Blog
有赞技术团队
有赞技术团队
小众软件
小众软件
IT之家
IT之家
Vercel News
Vercel News
V2EX - 技术
V2EX - 技术
雷峰网
雷峰网
L
Lohrmann on Cybersecurity
Cloudbric
Cloudbric
Engineering at Meta
Engineering at Meta
Schneier on Security
Schneier on Security
P
Privacy International News Feed
Apple Machine Learning Research
Apple Machine Learning Research
W
WeLiveSecurity
大猫的无限游戏
大猫的无限游戏
S
SegmentFault 最新的问题
J
Java Code Geeks
T
Threatpost
S
Secure Thoughts
T
Tailwind CSS Blog
V
V2EX
Attack and Defense Labs
Attack and Defense Labs
P
Palo Alto Networks Blog
S
Security @ Cisco Blogs
The GitHub Blog
The GitHub Blog
Simon Willison's Weblog
Simon Willison's Weblog
The Register - Security
The Register - Security
AWS News Blog
AWS News Blog
罗磊的独立博客
GbyAI
GbyAI
Blog — PlanetScale
Blog — PlanetScale
Microsoft Azure Blog
Microsoft Azure Blog
Forbes - Security
Forbes - Security
N
News | PayPal Newsroom
博客园 - 叶小钗
Hugging Face - Blog
Hugging Face - Blog
Exploit-DB.com RSS Feed
Exploit-DB.com RSS Feed
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻
Y
Y Combinator Blog
C
CXSECURITY Database RSS Feed - CXSecurity.com
Webroot Blog
Webroot Blog
爱范儿
爱范儿

DEV Community

Authentication Security Deep Dive: From Brute Force to Salted Hashing (With Java Examples) Why AI Systems Don’t Fail — They Drift Spilling beans for how i learn for exam😁"Reinforcement Learning Cheat Sheet" I Replaced Chrome with Safari for AI Browser Automation. Here's What Broke (and What Finally Worked) How Python Borrows Other People's Work The $40 Architecture: Processing 1 Billion API Requests with 99.99% Uptime Vibe Coding: A Workflow Guide (From Zero to SaaS) Most webhook security guides protect the wrong side. The scary part is delivery. Headless CMS for TanStack Start: Build a Blog with Cosmic EU Age Verification App "Hacked in 2 Minutes" — What Actually Happened Comfy Cloud’s delete function does not actually remove files Running AI Models on GPU Cloud Servers: A Beginner Guide Event-driven media intelligence with AWS Step Functions and Bedrock I scored 500 AI prompts across 8 quality dimensions — here's what broke How to Call Google Gemini API from Next.js (Free Tier, No Backend Needed) The Portal Protocol: Reclaiming Human Connection in the Age of AI How to Fix Your Team's Scattered Knowledge Problem With a Self-Hosted Forum Intro to tc Cloud Functors: A Graph-First Mental Model for the Modern Cloud Designing Multi-Tenant Backends With Both Ownership and Team Access I Built a Neumorphic CSS Library with 77+ Components — Here's What I Learned PostgreSQL Performance Optimization: Why Connection Pooling Is Critical at Scale Cómo construí un SaaS multi-rubro para gestionar expensas en Argentina con FastAPI + Vue 3 🚀 I Built an Ethical Hacking Scanner Tool – Open Source Project I Replaced /usage and /context in Claude Code With a Single Statusline A Pythonic Way to Handle Emails (IMAP/SMTP) with Auto-Discovery and AI-Ready Design I Collected 8.9 Million Polymarket Price Points — Here's What I Found About How Markets Really Move EcoTrack AI — Carbon Footprint Tracker & Dashboard Everyone's Using AI. No One Agrees How. 5 self-hosted ebook managers worth trying in 2026 Building Your First AI Agent with LangChain: From Chatbot to Autonomous Assistant Common SOC 2 Failures (Real World) Stop Vibe-Checking Your AI App: A Practical Guide to Evals How to Use SonarQube and SonarScanner Locally to Level Up Your Code Quality Your Next To-Do App Is Dead — I Replaced Mine with an OpenClaw AI Sign a Nostr event in 60 lines of Python using coincurve — no nostr-sdk, no nbxplorer, no rust toolchain ITGC Audit Explained Like You’re in Big 4 Patch Tuesday abril 2026: Microsoft parcha 163 vulnerabilidades y un zero-day en SharePoint Stop scraping everything: a better way to track competitor price changes Listing on MCPize + the Official MCP Registry while routing payments OUTSIDE the marketplace — how I kept 100% of my x402 revenue Building an AI-Powered Risk Intelligence System Using Serverless Architecture Why We Ripped Function Overloading Out of Our AI Toolchain Testing AI-Generated Code: How to Actually Know If It Works SaaS Churn Is Killing Your Business. Here Is What to Do About It (Without a Support Team) The Speed of AI Is No Longer Linear - And Self-Improving Models Are Why How to Implement RBAC for MCP Tools: A Practical Guide for Engineering Teams From Standard Quote to Persuasive Proposal: AI Automation for Arborists I built a CLI that scaffolds complete multi-tenant SaaS apps Axios CVE-2025–62718: The Silent SSRF Bug That Could Be Hiding in Your Node.js App Right Now The dashboard that ended our friendship Data Pipelines Explained Simply (and How to Build Them with Python) The Hidden Cost of AI Systems Nobody Talks About. undefined vs undeclared, and how typeof behaves Switching from file-based jobs to NATS/Kafka in Rust without changing code io_uring Adventures: Rust Servers That Love Syscalls Why Agentic AI is Killing the Traditional Database The POUR principles of web accessibility for developers and designers Quantum Neural Network 3D — A Deep Dive into Interactive WebGL Visualization How To Install Caveman In Codex On macOS And Windows Automation Pipeline Reliability: Why Your Workflow Breaks When Nobody Is Watching I Built an 'Open World' AI Coding Agent — It Works From ANY Folder From Freelancing to Product: A Tech Service Company's SaaS Transformation China's AI Giants: Adding Tencent Hunyuan & ByteDance Doubao to AI University (74 Providers) On the Vibe Coders and Their Lies clerk: Auto-Summarize Your Claude Code Sessions AI Weekly — 2026/04/10–04/17 | The Model Lockdown Is Here, but the Toolchain Is the Real Battleground AI 週報 — 2026/04/10–2026/04/17 模型封鎖潮來了,但工具鏈才是真戰場 Maybe this is how Open-Source apps are born... 🚀 Fine-Tune LLMs with LoRA and QLoRA: 2026 Guide tRPC v11 + Next.js App Router: End-to-End Type Safety Without the Boilerplate ShadCN UI in 2026: Why I Stopped Installing Component Libraries and Started Owning My Components SaaS Billing in React Server Components: Stripe + Supabase Without a Single `useEffect` Join our DEV Weekend Challenge — $1,000 in Prizes Across TEN winners! Submissions Due April 20 at 6:59 AM UTC. Implementing FSRS Spaced Repetition in Flutter + Supabase — Adding Memory Science to an AI Learning App "I Texted My Localhost From the Train — Claude Code Fixed the Bug Before I Got Home" I Built a Sales Prep AI and It Went Deeper Than Expected Design to Code #2: One JSON, Eleven Outputs Solving the 100M-Row Problem: A Summary Table Pattern for High-Volume Push Notification Logs Flutter Web With Wasm: What Actually Changes For Developers I Built 50 Royalty-Free Soundtracks for My Side Project in a Weekend Using AI Music Generation The Vibe Coding Security Checklist: 7 Things to Check Before You Ship Stop Letting Googlebot Guess Fix Your React App's SEO Right Desconstruindo o Streaming do LinkedIn: Como Criar um Engine de Extração de Vídeo de Alta Performance com HLS e FFmpeg (EDA Part-1) EDA (Exploratory Data Analysis) Explained With Real Life — Why Looking at Your Data Is the Most Important Step in Machine Learning Brand Relationship Management at Scale: Our 4-Touch Outreach System for 200+ Brands Why String.fromEnvironment() Might Return an Empty String in Dart JGuardrails 1.0.0 — Hardening Java LLM Apps Against Jailbreaks, Toxicity, and Prompt Injection Plan and Schedule a Full Week of Threads Content From One Claude Conversation Coding Cat Oran Ep3, Five Tables Changed Everything Updated: BFF Pattern I'm done watching freelancers get buried by 200 proposals. So I'm building the alternative. This is my first post BFS Algorithm in Java Step by Step Tutorial with Examples Tracking LLM Pricing Monthly: An Open Dataset for 22 AI Models How We Measure Content ROI on a Comparison Site: Revenue Attribution Without Perfect Data Introducing Nova AI Ops: The AI-Native Operating System for SRE Teams I built a free desktop video downloader for Windows — Grabbit How Talkie OCR Helps Vision-Impaired & Dyslexic Users Read the World Around Them VRCFaceTracking安装和iPhone面捕配置教程,有bug Even CrowdStrike Can't See Your Agents The Automation Gold Rush: What n8n Workflows and Claude Are Opening Up for Developers Right Now
How I Built an AI Document Ingestion Pipeline
Ryan Carter · 2026-04-29 · via DEV Community
<p><a href="https://github.com/meownoirsoft/symport" rel="noopener noreferrer">Symport</a> is an AI document ingestion pipeline that turns a phone photo of any paper document — receipt, EOB, prescription, utility bill — into structured JSON, then stores it in Postgres with embeddings for semantic search. The full flow is: image upload → Sharp preprocessing → GPT-4o vision extraction → normalized JSON → Postgres + pgvector. I built it because I hate paper and I also lose paper.</p> <p>This post walks through how the pipeline actually works, including the prompt engineering decisions that make extraction reliable enough to trust and the fallback layers that keep the app useful when extraction fails.</p> <h2> TL;DR </h2> <ul> <li> <strong>Stack:</strong> Sharp for image preprocessing, GPT-4o for vision extraction, Prisma + Postgres + pgvector for storage and semantic search.</li> <li> <strong>The extraction prompt does most of the work:</strong> explicit date context to fight year hallucinations, constrained <code>type</code>/<code>category</code> enums for predictable downstream branching, and a strict "JSON only, no markdown" tail.</li> <li> <strong>User correction loop:</strong> Users can add freeform feedback ("the drug name is metformin, not metFORMIN") and re-run extraction; the feedback gets injected back into the system prompt.</li> <li> <strong>Schema choice:</strong> A single <code>extractedData</code> JSON column instead of per-type tables, with a denormalized <code>searchText</code> field for fast keyword search and an <code>embedding</code> column for semantic search.</li> <li> <strong>Two fallback layers:</strong> Document still saves if there's no API key, and still saves with an error summary if extraction throws — nothing is ever lost because AI had a bad day.</li> </ul> <h2> What the pipeline does </h2> <p>The flow is straightforward:<br> </p> <div class="highlight js-code-highlight"> <pre class="highlight plaintext"><code>Image upload → sharpen + encode → GPT-4o vision → structured JSON → Postgres + embeddings </code></pre> </div> <p>A user photographs a receipt, an insurance EOB, a prescription, a utility bill — anything on paper. The app returns a structured JSON object with the relevant fields extracted, tagged, and ready to query. No manual data entry.</p> <h2> Step 1: Image preprocessing </h2> <p>Raw phone photos are large and often noisy. Before sending to the vision model, every image gets sharpened and re-encoded using Sharp:<br> </p> <div class="highlight js-code-highlight"> <pre class="highlight typescript"><code><span class="kd">const</span> <span class="nx">rawBuffer</span> <span class="o">=</span> <span class="nx">Buffer</span><span class="p">.</span><span class="k">from</span><span class="p">(</span><span class="k">await</span> <span class="nx">file</span><span class="p">.</span><span class="nf">arrayBuffer</span><span class="p">());</span> <span class="kd">const</span> <span class="nx">buffer</span> <span class="o">=</span> <span class="k">await</span> <span class="nf">sharpenAndEncode</span><span class="p">(</span><span class="nx">rawBuffer</span><span class="p">);</span> </code></pre> </div> <p>Sharp handles resizing, sharpening, and JPEG re-encoding in one pass. This serves two purposes: it reduces the payload size for the API call, and sharpening improves OCR accuracy on text-heavy documents like receipts. A blurry photo of small print is genuinely harder for vision models — a little preprocessing pays off.</p> <p>The processed image gets saved to disk as the source of truth, then the buffer goes to the extraction pipeline:<br> </p> <div class="highlight js-code-highlight"> <pre class="highlight typescript"><code><span class="kd">const</span> <span class="nx">filename</span> <span class="o">=</span> <span class="s2">`</span><span class="p">${</span><span class="nf">randomBytes</span><span class="p">(</span><span class="mi">12</span><span class="p">).</span><span class="nf">toString</span><span class="p">(</span><span class="dl">"</span><span class="s2">hex</span><span class="dl">"</span><span class="p">)}</span><span class="s2">.jpg`</span><span class="p">;</span> <span class="k">await</span> <span class="nf">writeFile</span><span class="p">(</span><span class="nx">fullPath</span><span class="p">,</span> <span class="nx">buffer</span><span class="p">);</span> <span class="kd">const</span> <span class="nx">extracted</span> <span class="o">=</span> <span class="k">await</span> <span class="nf">extractFromImageBuffer</span><span class="p">(</span><span class="nx">buffer</span><span class="p">);</span> </code></pre> </div> <p>Random hex filename prevents collisions and avoids leaking any metadata about the document in the path.</p> <h2> Step 2: The extraction prompt </h2> <p>This is where most of the real engineering lives. The system prompt does a lot of work to make the model's output consistent and parseable.</p> <p>The prompt has three parts assembled at startup:<br> </p> <div class="highlight js-code-highlight"> <pre class="highlight typescript"><code><span class="kd">const</span> <span class="nx">EXTRACTION_SYSTEM_HEAD</span> <span class="o">=</span> <span class="s2">`You are a document extraction assistant. Analyze the image and extract structured data. Current date context: We are in 2025. Use 2025 (not 2023 or other past years) for any ambiguous or partial dates when no stronger clue is present. Use context clues from the document text to infer the correct year: - "2025 taxes due in 2026" → tax year 2025 - "Plan year 2025", "Coverage year 2025" → use 2025 - "Due in 2026" on a tax-related doc often refers to tax year 2025 Respond with a single JSON object. Include "type", "category", "title", and "tags" in every response. - "type": one of rx_receipt, eob, utility_bill, general - "category": one of receipt, financial, medical, government, legal, identity, general - "title" (required, 2-5 words max): Short label only. No sentences. - "tags": array of 3–8 short labels. No spaces; use underscores if needed. `</span><span class="p">;</span> </code></pre> </div> <p>A few decisions worth calling out here:</p> <p><strong>Explicit date context.</strong> Vision models can hallucinate dates, especially on documents where the year is ambiguous. Anchoring the prompt with the current year and showing examples of how to reason about year context dramatically reduces date errors. Without this, a 2025 tax document might come back with 2023 dates because the model defaulted to its training data.</p> <p><strong>Constrained type and category values.</strong> Giving the model an explicit enum for <code>type</code> and <code>category</code> means you get predictable values you can branch on in code. Open-ended classification produces inconsistent strings that are annoying to handle downstream.</p> <p><strong>Short title constraint.</strong> "2-5 words max, no sentences" prevents the model from writing a summary disguised as a title. You want "Prescription receipt" not "This document appears to be a receipt from Walgreens for a prescription medication."</p> <p>The tail of the prompt closes with:<br> </p> <div class="highlight js-code-highlight"> <pre class="highlight typescript"><code><span class="kd">const</span> <span class="nx">EXTRACTION_SYSTEM_TAIL</span> <span class="o">=</span> <span class="s2">` Use null for missing values. Amounts as numbers. Dates as YYYY-MM-DD; use context clues for year. Output only valid JSON, no markdown or explanation.`</span><span class="p">;</span> </code></pre> </div> <p>"Output only valid JSON, no markdown or explanation" is load-bearing. Without it, GPT-4o will frequently wrap the response in a markdown code block. The extraction code handles that case anyway, but telling the model not to do it reduces the cleanup work:<br> </p> <div class="highlight js-code-highlight"> <pre class="highlight typescript"><code><span class="c1">// Strip optional markdown code block</span> <span class="kd">let</span> <span class="nx">jsonStr</span> <span class="o">=</span> <span class="nx">raw</span><span class="p">;</span> <span class="kd">const</span> <span class="nx">match</span> <span class="o">=</span> <span class="nx">raw</span><span class="p">.</span><span class="nf">match</span><span class="p">(</span><span class="sr">/``</span><span class="err">` </span><span class="p">{</span><span class="o">%</span> <span class="nx">endraw</span> <span class="o">%</span><span class="p">}</span> <span class="p">(?:</span><span class="nx">json</span><span class="p">)?</span><span class="err">\</span><span class="nx">s</span><span class="o">*</span><span class="p">([</span><span class="err">\</span><span class="nx">s</span><span class="err">\</span><span class="nx">S</span><span class="p">]</span><span class="o">*</span><span class="p">?)</span> <span class="p">{</span><span class="o">%</span> <span class="nx">raw</span> <span class="o">%</span><span class="p">}</span> <span class="s2">```/); if (match) jsonStr = match[1].trim(); </span></code></pre> </div> <h2> Step 3: User feedback loop </h2> <p>One of the more useful features is the ability to correct extractions. If the model gets something wrong — misreads a drug name, gets the date wrong, miscategorizes the document — the user can add a correction note and re-run extraction. That feedback gets injected directly into the system prompt:<br> </p> <div class="highlight js-code-highlight"> <pre class="highlight typescript"><code><span class="k">if </span><span class="p">(</span><span class="nx">options</span><span class="p">?.</span><span class="nx">userFeedback</span><span class="p">?.</span><span class="nf">trim</span><span class="p">())</span> <span class="p">{</span> <span class="nx">systemContent</span> <span class="o">+=</span> <span class="s2">`\n\nIMPORTANT - User feedback on this document (apply these corrections): </span><span class="p">${</span><span class="nx">options</span><span class="p">.</span><span class="nx">userFeedback</span><span class="p">.</span><span class="nf">trim</span><span class="p">()}</span><span class="s2">`</span><span class="p">;</span> <span class="p">}</span> </code></pre> </div> <p>This means the model gets a second pass with explicit correction instructions. In practice it works well — "the drug name is metformin not metFORMIN" or "this is a 2025 EOB not 2024" gets applied reliably.</p> <p>The feedback also gets stored in the database as <code>extractionNotes</code> on the document, so you have a record of what was corrected.</p> <h2> Step 4: The data model </h2> <p>The Prisma schema keeps things straightforward:<br> </p> <div class="highlight js-code-highlight"> <pre class="highlight typescript"><code><span class="nx">model</span> <span class="nx">Document</span> <span class="p">{</span> <span class="nx">id</span> <span class="nb">String</span> <span class="p">@</span><span class="nd">id</span> <span class="p">@</span><span class="nd">default</span><span class="p">(</span><span class="nf">cuid</span><span class="p">())</span> <span class="nx">imagePath</span> <span class="nb">String</span><span class="p">?</span> <span class="nx">noteText</span> <span class="nb">String</span><span class="p">?</span> <span class="nx">status</span> <span class="nb">String</span> <span class="p">@</span><span class="nd">default</span><span class="p">(</span><span class="dl">"</span><span class="s2">pending</span><span class="dl">"</span><span class="p">)</span> <span class="nx">extractedData</span> <span class="nx">Json</span> <span class="nx">searchText</span> <span class="nb">String</span><span class="p">?</span> <span class="nx">embedding</span> <span class="nc">Unsupported</span><span class="p">(</span><span class="dl">"</span><span class="s2">vector(1536)</span><span class="dl">"</span><span class="p">)?</span> <span class="nx">tags</span> <span class="nb">String</span><span class="p">[]</span> <span class="nx">extractionNotes</span> <span class="nb">String</span><span class="p">?</span> <span class="nx">createdAt</span> <span class="nx">DateTime</span> <span class="p">@</span><span class="nd">default</span><span class="p">(</span><span class="nf">now</span><span class="p">())</span> <span class="nx">updatedAt</span> <span class="nx">DateTime</span> <span class="p">@</span><span class="nd">updatedAt</span> <span class="p">}</span> </code></pre> </div> <p>A few design choices here:</p> <p><strong><code>extractedData</code> is a JSON blob.</strong> Rather than creating separate tables for each document type (receipts, EOBs, utility bills), all extracted data lives in a single JSON column. This makes the schema flexible — different document types have different fields, and a rigid relational schema would be a constant maintenance burden as new types are added.</p> <p><strong><code>searchText</code> is denormalized.</strong> After extraction, key fields get pulled out and concatenated into a single <code>searchText</code> string for full-text search. This is faster to query than parsing JSON at search time:<br> </p> <div class="highlight js-code-highlight"> <pre class="highlight typescript"><code><span class="k">export</span> <span class="kd">function</span> <span class="nf">buildSearchText</span><span class="p">(</span><span class="nx">data</span><span class="p">:</span> <span class="nx">ExtractedDoc</span><span class="p">):</span> <span class="kr">string</span> <span class="p">{</span> <span class="kd">const</span> <span class="nx">parts</span><span class="p">:</span> <span class="kr">string</span><span class="p">[]</span> <span class="o">=</span> <span class="p">[];</span> <span class="nx">parts</span><span class="p">.</span><span class="nf">push</span><span class="p">(</span><span class="nf">effectiveTitle</span><span class="p">(</span><span class="nx">data</span><span class="p">));</span> <span class="k">if </span><span class="p">(</span><span class="dl">"</span><span class="s2">summary</span><span class="dl">"</span> <span class="k">in</span> <span class="nx">data</span> <span class="o">&amp;&amp;</span> <span class="nx">data</span><span class="p">.</span><span class="nx">summary</span><span class="p">)</span> <span class="nx">parts</span><span class="p">.</span><span class="nf">push</span><span class="p">(</span><span class="nc">String</span><span class="p">(</span><span class="nx">data</span><span class="p">.</span><span class="nx">summary</span><span class="p">));</span> <span class="k">if </span><span class="p">(</span><span class="dl">"</span><span class="s2">drug_name</span><span class="dl">"</span> <span class="k">in</span> <span class="nx">data</span> <span class="o">&amp;&amp;</span> <span class="nx">data</span><span class="p">.</span><span class="nx">drug_name</span><span class="p">)</span> <span class="nx">parts</span><span class="p">.</span><span class="nf">push</span><span class="p">(</span><span class="nc">String</span><span class="p">(</span><span class="nx">data</span><span class="p">.</span><span class="nx">drug_name</span><span class="p">));</span> <span class="k">if </span><span class="p">(</span><span class="dl">"</span><span class="s2">insurer</span><span class="dl">"</span> <span class="k">in</span> <span class="nx">data</span> <span class="o">&amp;&amp;</span> <span class="nx">data</span><span class="p">.</span><span class="nx">insurer</span><span class="p">)</span> <span class="nx">parts</span><span class="p">.</span><span class="nf">push</span><span class="p">(</span><span class="nc">String</span><span class="p">(</span><span class="nx">data</span><span class="p">.</span><span class="nx">insurer</span><span class="p">));</span> <span class="k">if </span><span class="p">(</span><span class="dl">"</span><span class="s2">tags</span><span class="dl">"</span> <span class="k">in</span> <span class="nx">data</span> <span class="o">&amp;&amp;</span> <span class="nb">Array</span><span class="p">.</span><span class="nf">isArray</span><span class="p">(</span><span class="nx">data</span><span class="p">.</span><span class="nx">tags</span><span class="p">))</span> <span class="p">{</span> <span class="nx">parts</span><span class="p">.</span><span class="nf">push</span><span class="p">(...</span><span class="nx">data</span><span class="p">.</span><span class="nx">tags</span><span class="p">.</span><span class="nf">map</span><span class="p">(</span><span class="nx">t</span> <span class="o">=&gt;</span> <span class="nc">String</span><span class="p">(</span><span class="nx">t</span><span class="p">).</span><span class="nf">trim</span><span class="p">()).</span><span class="nf">filter</span><span class="p">(</span><span class="nb">Boolean</span><span class="p">));</span> <span class="p">}</span> <span class="k">return</span> <span class="nx">parts</span><span class="p">.</span><span class="nf">join</span><span class="p">(</span><span class="dl">"</span><span class="s2"> </span><span class="dl">"</span><span class="p">);</span> <span class="p">}</span> </code></pre> </div> <p><strong><code>embedding</code> for semantic search.</strong> After the document is saved, an embedding gets generated from <code>searchText</code> and stored in a pgvector column. This enables semantic search — finding "cholesterol medication" when the document says "lipitor" — without a separate vector database. Just pgvector as a Postgres extension.</p> <h2> Step 5: Graceful degradation </h2> <p>The pipeline has two fallback layers. First, if there's no API key configured, the document still gets saved — just without extraction:<br> </p> <div class="highlight js-code-highlight"> <pre class="highlight typescript"><code><span class="k">if </span><span class="p">(</span><span class="o">!</span><span class="nx">process</span><span class="p">.</span><span class="nx">env</span><span class="p">.</span><span class="nx">OPENAI_API_KEY</span><span class="p">)</span> <span class="p">{</span> <span class="k">await</span> <span class="nx">prisma</span><span class="p">.</span><span class="nb">document</span><span class="p">.</span><span class="nf">create</span><span class="p">({</span> <span class="na">data</span><span class="p">:</span> <span class="p">{</span> <span class="na">imagePath</span><span class="p">:</span> <span class="nx">filename</span><span class="p">,</span> <span class="na">status</span><span class="p">:</span> <span class="dl">"</span><span class="s2">pending</span><span class="dl">"</span><span class="p">,</span> <span class="na">extractedData</span><span class="p">:</span> <span class="p">{</span> <span class="na">type</span><span class="p">:</span> <span class="dl">"</span><span class="s2">general</span><span class="dl">"</span><span class="p">,</span> <span class="na">title</span><span class="p">:</span> <span class="dl">"</span><span class="s2">Document</span><span class="dl">"</span><span class="p">,</span> <span class="na">summary</span><span class="p">:</span> <span class="dl">"</span><span class="s2">Extraction skipped (no OPENAI_API_KEY)</span><span class="dl">"</span> <span class="p">},</span> <span class="na">searchText</span><span class="p">:</span> <span class="dl">"</span><span class="s2">Extraction skipped</span><span class="dl">"</span><span class="p">,</span> <span class="na">tags</span><span class="p">:</span> <span class="p">[],</span> <span class="p">},</span> <span class="p">});</span> <span class="p">}</span> </code></pre> </div> <p>Second, if extraction throws, the document still gets saved with an error summary rather than failing the whole request:<br> </p> <div class="highlight js-code-highlight"> <pre class="highlight typescript"><code><span class="k">try</span> <span class="p">{</span> <span class="kd">const</span> <span class="nx">extracted</span> <span class="o">=</span> <span class="k">await</span> <span class="nf">extractFromImageBuffer</span><span class="p">(</span><span class="nx">buffer</span><span class="p">);</span> <span class="nx">extractedData</span> <span class="o">=</span> <span class="nx">extracted</span> <span class="k">as</span> <span class="nb">Record</span><span class="o">&lt;</span><span class="kr">string</span><span class="p">,</span> <span class="nx">unknown</span><span class="o">&gt;</span><span class="p">;</span> <span class="p">}</span> <span class="k">catch </span><span class="p">(</span><span class="nx">err</span><span class="p">)</span> <span class="p">{</span> <span class="nx">extractedData</span> <span class="o">=</span> <span class="p">{</span> <span class="na">type</span><span class="p">:</span> <span class="dl">"</span><span class="s2">general</span><span class="dl">"</span><span class="p">,</span> <span class="na">summary</span><span class="p">:</span> <span class="dl">"</span><span class="s2">Extraction failed: </span><span class="dl">"</span> <span class="o">+</span> <span class="p">(</span><span class="nx">err</span> <span class="k">instanceof</span> <span class="nb">Error</span> <span class="p">?</span> <span class="nx">err</span><span class="p">.</span><span class="nx">message</span> <span class="p">:</span> <span class="dl">"</span><span class="s2">Unknown error</span><span class="dl">"</span><span class="p">),</span> <span class="p">};</span> <span class="p">}</span> </code></pre> </div> <p>The image is always saved. The extraction is best-effort. Users can re-trigger extraction manually, or add correction notes and re-run. Nothing gets lost because AI had a bad day.</p> <h2> The model </h2> <p>The extraction model is configurable via environment variable with <code>gpt-4o</code> as the default:<br> </p> <div class="highlight js-code-highlight"> <pre class="highlight typescript"><code><span class="kd">const</span> <span class="nx">model</span> <span class="o">=</span> <span class="nx">process</span><span class="p">.</span><span class="nx">env</span><span class="p">.</span><span class="nx">OPENAI_EXTRACTION_MODEL</span> <span class="o">||</span> <span class="dl">"</span><span class="s2">gpt-4o</span><span class="dl">"</span><span class="p">;</span> </code></pre> </div> <p>GPT-4o is the right choice here — it's genuinely better than smaller models at reading degraded document images, handwriting, and small print. For this specific task the quality difference is noticeable enough to justify the cost. Document extraction is a write-time operation (not a search-time one), so the latency and cost are acceptable.</p> <h2> What I'd do differently </h2> <p>A few things I'd change with hindsight:</p> <p><strong>Add a confidence score.</strong> The model sometimes hedges on fields it's uncertain about — a low-confidence flag on individual fields would let the UI highlight things that need user review rather than silently storing potentially wrong data.</p> <p><strong>Chunk large documents.</strong> A single-page receipt is fine. A multi-page insurance EOB or medical record is harder — the model gets less accurate as documents get longer or more complex. Chunking multi-page documents and merging the extracted JSON would improve accuracy on longer content.</p> <p><strong>Store the raw extraction response.</strong> Right now only the normalized result gets stored. Keeping the raw model output alongside it would make debugging extraction issues much easier.</p> <p>The full source is on GitHub at <a href="https://github.com/meownoirsoft/symport" rel="noopener noreferrer">github.com/meownoirsoft/symport</a>. The extraction logic lives in <code>lib/extract.ts</code> and the ingestion endpoint is <code>app/api/documents/route.ts</code> if you want to dig in.</p> <h2> FAQ </h2> <h3> Why GPT-4o for vision instead of a cheaper model or open-source alternative? </h3> <p>GPT-4o reads degraded phone photos, handwriting, and small print noticeably better than smaller or open-source vision models. For document extraction, getting the dates and amounts wrong is a much bigger problem than the per-call cost, so paying for the better model is worth it. Extraction runs once at write time, not on every read.</p> <h3> How do you stop the model from hallucinating dates or amounts? </h3> <p>The biggest wins are anchoring "current year" context in the system prompt with explicit examples, asking the model to use context clues from the document itself ("plan year 2025", "due in 2026" → tax year 2025), and constraining types/categories to enums so the model can't drift. The user-feedback loop catches anything that still slips through.</p> <h3> Why store extracted data as a single JSON column instead of typed tables? </h3> <p>Different document types have different fields — a prescription receipt has <code>drug_name</code> and <code>pharmacy</code>, an EOB has <code>insurer</code> and <code>claim_id</code>, a utility bill has <code>account_number</code> and <code>service_period</code>. A relational schema for every variant would be a constant migration treadmill. JSON keeps the schema flexible, and the denormalized <code>searchText</code> and <code>embedding</code> columns make queries fast where it matters.</p> <h3> What happens if the model returns invalid JSON? </h3> <p>The extraction code strips optional markdown code fences (<br> <br> <code>json ...</code><br> <br> ) and parses the rest. If parsing still fails, the document saves with an error summary in the <code>extractedData.summary</code> field rather than throwing — the user can re-run extraction or add a correction note. The image and metadata are never lost.</p> <h3> Can I use this with Anthropic's Claude vision instead of GPT-4o? </h3> <p>Yes. The extraction prompt is provider-agnostic and the model is configurable via <code>OPENAI_EXTRACTION_MODEL</code>. Swap the SDK call for the Anthropic SDK (or route through OpenRouter to avoid a code change) and Claude's vision models work as a drop-in alternative. The "JSON only, no markdown" instruction is even more important on Claude — it likes to explain itself by default.</p> <h3> How do you handle multi-page documents? </h3> <p>Today the pipeline treats each photo as a single document, which is fine for single-page items (receipts, prescriptions). For multi-page EOBs or medical records, the right next step is to chunk the document into pages, run extraction per page, and merge the resulting JSON into a single record. Adding that is on the to-do list.</p>