惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

云风的 BLOG
云风的 BLOG
Blog — PlanetScale
Blog — PlanetScale
博客园 - 【当耐特】
博客园_首页
The GitHub Blog
The GitHub Blog
月光博客
月光博客
Hugging Face - Blog
Hugging Face - Blog
有赞技术团队
有赞技术团队
博客园 - 三生石上(FineUI控件)
D
Docker
Stack Overflow Blog
Stack Overflow Blog
WordPress大学
WordPress大学
钛媒体:引领未来商业与生活新知
钛媒体:引领未来商业与生活新知
Apple Machine Learning Research
Apple Machine Learning Research
Vercel News
Vercel News
酷 壳 – CoolShell
酷 壳 – CoolShell
雷峰网
雷峰网
小众软件
小众软件
I
InfoQ
A
About on SuperTechFans
T
The Blog of Author Tim Ferriss
S
SegmentFault 最新的问题
Microsoft Azure Blog
Microsoft Azure Blog
博客园 - Franky

DEV Community

Authentication Security Deep Dive: From Brute Force to Salted Hashing (With Java Examples) Why AI Systems Don’t Fail — They Drift Spilling beans for how i learn for exam😁"Reinforcement Learning Cheat Sheet" I Replaced Chrome with Safari for AI Browser Automation. Here's What Broke (and What Finally Worked) How Python Borrows Other People's Work The $40 Architecture: Processing 1 Billion API Requests with 99.99% Uptime Vibe Coding: A Workflow Guide (From Zero to SaaS) Most webhook security guides protect the wrong side. The scary part is delivery. Headless CMS for TanStack Start: Build a Blog with Cosmic EU Age Verification App "Hacked in 2 Minutes" — What Actually Happened Comfy Cloud’s delete function does not actually remove files Running AI Models on GPU Cloud Servers: A Beginner Guide Event-driven media intelligence with AWS Step Functions and Bedrock I scored 500 AI prompts across 8 quality dimensions — here's what broke How to Call Google Gemini API from Next.js (Free Tier, No Backend Needed) The Portal Protocol: Reclaiming Human Connection in the Age of AI How to Fix Your Team's Scattered Knowledge Problem With a Self-Hosted Forum Intro to tc Cloud Functors: A Graph-First Mental Model for the Modern Cloud Designing Multi-Tenant Backends With Both Ownership and Team Access I Built a Neumorphic CSS Library with 77+ Components — Here's What I Learned PostgreSQL Performance Optimization: Why Connection Pooling Is Critical at Scale Cómo construí un SaaS multi-rubro para gestionar expensas en Argentina con FastAPI + Vue 3 🚀 I Built an Ethical Hacking Scanner Tool – Open Source Project I Replaced /usage and /context in Claude Code With a Single Statusline A Pythonic Way to Handle Emails (IMAP/SMTP) with Auto-Discovery and AI-Ready Design I Collected 8.9 Million Polymarket Price Points — Here's What I Found About How Markets Really Move EcoTrack AI — Carbon Footprint Tracker & Dashboard Everyone's Using AI. No One Agrees How. 5 self-hosted ebook managers worth trying in 2026 Building Your First AI Agent with LangChain: From Chatbot to Autonomous Assistant
I Tested 15 LLMs for Web Scraping and Built Heuristics In...
Rohith M · 2026-05-06 · via DEV Community

Rohith M

Clura AI web scraper Chrome extension detecting fields on a business directory website

The problem nobody talks about: 600KB of DOM

When I started building a web scraper, the obvious move was to send the page to an LLM and ask it to extract the data. Simple, right?

Wrong. A typical product listing page is 500–700KB of raw DOM. Sending that to any model means you're paying for ~150,000 tokens per page, waiting 15–30 seconds per request, and hitting context limits on anything complex.

I hit this wall on page one.

Four months, 15 models, same result

I tested everything: GPT-4, GPT-4o, Gemini 1.5 Pro, Gemini Ultra, Claude 3 Opus, Claude 3.5 Sonnet, Mistral Large, Llama 3 70B, Cohere Command R+, and a handful of smaller fine-tuned models.

The results were consistent:

  • GPT-4 / Gemini Ultra: accurate, but 25–35 seconds per page
  • Claude 3.5 Sonnet: best accuracy-to-latency ratio, still 5–10 seconds
  • Smaller models: faster, but hallucinated field names constantly

No model solved the latency problem because I was asking them to solve the wrong problem.

The pre-processor breakthrough

The real issue wasn't the model — it was the input size.

I built a DOM pre-processor:

  1. Strip all <script>, <style>, and tracking pixels
  2. Remove navigation, footer, sidebar elements
  3. Collapse deeply nested wrappers that carry no semantic content
  4. Apply SimHash to deduplicate structurally identical subtrees

Result: 580KB → 4.2KB. A 99.3% reduction.

With a 4KB input, every model became fast. But something more interesting happened: at that size, the repeating patterns became obvious. The same structure repeated 20, 50, 100 times — product cards, directory rows, search results.

How AI web scraping works step by step

The architecture decision

If the pattern is already obvious from the structure alone, why am I paying a model to find it?

I wrote a heuristic detector:

  • Identify elements with 3+ structurally identical siblings
  • Score candidate lists by depth, child count uniformity, and text density
  • Return ranked list candidates in 0ms

Then AI enters after detection — not to identify the list, but to label the fields and structure the output. That's a 200-token job, not a 150,000-token job.

Step Approach Latency
List detection Heuristics 0.2ms
Field labeling LLM (small input) ~2s
Total ~2s

vs. naive LLM approach: 25–35 seconds.

What I actually shipped

This architecture became Clura — a heuristic-first AI web scraper Chrome extension.

Open any page, Clura automatically detects every list using the heuristic engine. You pick the list, pick the fields you want, and it extracts all records in seconds. No "describe what you want" prompts. No robot training. No 30-second waits. The heuristic layer handles detection; AI handles labeling.

The lesson

LLMs are exceptional at understanding what something means. They're terrible at scanning 600KB of HTML to find where something is. That's a structural pattern problem — and structural pattern problems are what algorithms are for.

The best AI product architecture I've found isn't "use the best model." It's "use heuristics to reduce the problem until the model only sees what it's actually good at."

If you're building anything with LLMs on messy real-world inputs, the DOM pre-processing step alone is worth stealing. It will make every model you use faster, cheaper, and more accurate — regardless of the underlying task.


If you want to see this in action, try Clura free — it runs entirely in your browser with no server round-trips for detection.