惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

P
Proofpoint News Feed
H
Hacker News: Front Page
C
CXSECURITY Database RSS Feed - CXSecurity.com
C
Cisco Blogs
P
Palo Alto Networks Blog
Know Your Adversary
Know Your Adversary
D
Darknet – Hacking Tools, Hacker News & Cyber Security
C
Cybersecurity and Infrastructure Security Agency CISA
AWS News Blog
AWS News Blog
Spread Privacy
Spread Privacy
S
Schneier on Security
The Hacker News
The Hacker News
Cyberwarzone
Cyberwarzone
T
Tenable Blog
C
Cyber Attacks, Cyber Crime and Cyber Security
K
KPMG report finds enterprise disconnect between AI and its ROI | CIO
T
Tailwind CSS Blog
S
Secure Thoughts
N
Netflix TechBlog - Medium
T
The Exploit Database - CXSecurity.com
I
Intezer
Application and Cybersecurity Blog
Application and Cybersecurity Blog
Help Net Security
Help Net Security
K
Kaspersky official blog
Google Online Security Blog
Google Online Security Blog
L
LangChain Blog
Martin Fowler
Martin Fowler
L
LINUX DO - 热门话题
Hacker News: Ask HN
Hacker News: Ask HN
www.infosecurity-magazine.com
www.infosecurity-magazine.com
有赞技术团队
有赞技术团队
P
Privacy International News Feed
cs.CV updates on arXiv.org
cs.CV updates on arXiv.org
Recent Announcements
Recent Announcements
cs.CL updates on arXiv.org
cs.CL updates on arXiv.org
cs.AI updates on arXiv.org
cs.AI updates on arXiv.org
让小产品的独立变现更简单 - ezindie.com
让小产品的独立变现更简单 - ezindie.com
The Register - Security
The Register - Security
云风的 BLOG
云风的 BLOG
Google DeepMind News
Google DeepMind News
阮一峰的网络日志
阮一峰的网络日志
WordPress大学
WordPress大学
Recorded Future
Recorded Future
The Last Watchdog
The Last Watchdog
G
Google Developers Blog
T
Threatpost
小众软件
小众软件
S
Securelist
Recent Commits to openclaw:main
Recent Commits to openclaw:main
O
OpenAI News

DEV Community

Authentication Security Deep Dive: From Brute Force to Salted Hashing (With Java Examples) Why AI Systems Don’t Fail — They Drift Spilling beans for how i learn for exam😁"Reinforcement Learning Cheat Sheet" I Replaced Chrome with Safari for AI Browser Automation. Here's What Broke (and What Finally Worked) How Python Borrows Other People's Work The $40 Architecture: Processing 1 Billion API Requests with 99.99% Uptime Vibe Coding: A Workflow Guide (From Zero to SaaS) Most webhook security guides protect the wrong side. The scary part is delivery. Headless CMS for TanStack Start: Build a Blog with Cosmic EU Age Verification App "Hacked in 2 Minutes" — What Actually Happened Comfy Cloud’s delete function does not actually remove files Running AI Models on GPU Cloud Servers: A Beginner Guide Event-driven media intelligence with AWS Step Functions and Bedrock I scored 500 AI prompts across 8 quality dimensions — here's what broke How to Call Google Gemini API from Next.js (Free Tier, No Backend Needed) The Portal Protocol: Reclaiming Human Connection in the Age of AI How to Fix Your Team's Scattered Knowledge Problem With a Self-Hosted Forum Intro to tc Cloud Functors: A Graph-First Mental Model for the Modern Cloud Designing Multi-Tenant Backends With Both Ownership and Team Access I Built a Neumorphic CSS Library with 77+ Components — Here's What I Learned PostgreSQL Performance Optimization: Why Connection Pooling Is Critical at Scale Cómo construí un SaaS multi-rubro para gestionar expensas en Argentina con FastAPI + Vue 3 🚀 I Built an Ethical Hacking Scanner Tool – Open Source Project I Replaced /usage and /context in Claude Code With a Single Statusline A Pythonic Way to Handle Emails (IMAP/SMTP) with Auto-Discovery and AI-Ready Design I Collected 8.9 Million Polymarket Price Points — Here's What I Found About How Markets Really Move EcoTrack AI — Carbon Footprint Tracker & Dashboard Everyone's Using AI. No One Agrees How. 5 self-hosted ebook managers worth trying in 2026 Building Your First AI Agent with LangChain: From Chatbot to Autonomous Assistant Common SOC 2 Failures (Real World) Stop Vibe-Checking Your AI App: A Practical Guide to Evals How to Use SonarQube and SonarScanner Locally to Level Up Your Code Quality Your Next To-Do App Is Dead — I Replaced Mine with an OpenClaw AI Sign a Nostr event in 60 lines of Python using coincurve — no nostr-sdk, no nbxplorer, no rust toolchain ITGC Audit Explained Like You’re in Big 4 Patch Tuesday abril 2026: Microsoft parcha 163 vulnerabilidades y un zero-day en SharePoint Stop scraping everything: a better way to track competitor price changes Listing on MCPize + the Official MCP Registry while routing payments OUTSIDE the marketplace — how I kept 100% of my x402 revenue Building an AI-Powered Risk Intelligence System Using Serverless Architecture Why We Ripped Function Overloading Out of Our AI Toolchain Testing AI-Generated Code: How to Actually Know If It Works SaaS Churn Is Killing Your Business. Here Is What to Do About It (Without a Support Team) The Speed of AI Is No Longer Linear - And Self-Improving Models Are Why How to Implement RBAC for MCP Tools: A Practical Guide for Engineering Teams From Standard Quote to Persuasive Proposal: AI Automation for Arborists I built a CLI that scaffolds complete multi-tenant SaaS apps Axios CVE-2025–62718: The Silent SSRF Bug That Could Be Hiding in Your Node.js App Right Now The dashboard that ended our friendship Data Pipelines Explained Simply (and How to Build Them with Python) The Hidden Cost of AI Systems Nobody Talks About. undefined vs undeclared, and how typeof behaves Switching from file-based jobs to NATS/Kafka in Rust without changing code io_uring Adventures: Rust Servers That Love Syscalls Why Agentic AI is Killing the Traditional Database The POUR principles of web accessibility for developers and designers Quantum Neural Network 3D — A Deep Dive into Interactive WebGL Visualization How To Install Caveman In Codex On macOS And Windows Automation Pipeline Reliability: Why Your Workflow Breaks When Nobody Is Watching I Built an 'Open World' AI Coding Agent — It Works From ANY Folder From Freelancing to Product: A Tech Service Company's SaaS Transformation China's AI Giants: Adding Tencent Hunyuan & ByteDance Doubao to AI University (74 Providers) On the Vibe Coders and Their Lies clerk: Auto-Summarize Your Claude Code Sessions AI Weekly — 2026/04/10–04/17 | The Model Lockdown Is Here, but the Toolchain Is the Real Battleground AI 週報 — 2026/04/10–2026/04/17 模型封鎖潮來了,但工具鏈才是真戰場 Maybe this is how Open-Source apps are born... 🚀 Fine-Tune LLMs with LoRA and QLoRA: 2026 Guide tRPC v11 + Next.js App Router: End-to-End Type Safety Without the Boilerplate ShadCN UI in 2026: Why I Stopped Installing Component Libraries and Started Owning My Components SaaS Billing in React Server Components: Stripe + Supabase Without a Single `useEffect` Join our DEV Weekend Challenge — $1,000 in Prizes Across TEN winners! Submissions Due April 20 at 6:59 AM UTC. Implementing FSRS Spaced Repetition in Flutter + Supabase — Adding Memory Science to an AI Learning App "I Texted My Localhost From the Train — Claude Code Fixed the Bug Before I Got Home" I Built a Sales Prep AI and It Went Deeper Than Expected Design to Code #2: One JSON, Eleven Outputs Solving the 100M-Row Problem: A Summary Table Pattern for High-Volume Push Notification Logs Flutter Web With Wasm: What Actually Changes For Developers I Built 50 Royalty-Free Soundtracks for My Side Project in a Weekend Using AI Music Generation The Vibe Coding Security Checklist: 7 Things to Check Before You Ship Stop Letting Googlebot Guess Fix Your React App's SEO Right Desconstruindo o Streaming do LinkedIn: Como Criar um Engine de Extração de Vídeo de Alta Performance com HLS e FFmpeg (EDA Part-1) EDA (Exploratory Data Analysis) Explained With Real Life — Why Looking at Your Data Is the Most Important Step in Machine Learning Brand Relationship Management at Scale: Our 4-Touch Outreach System for 200+ Brands Why String.fromEnvironment() Might Return an Empty String in Dart JGuardrails 1.0.0 — Hardening Java LLM Apps Against Jailbreaks, Toxicity, and Prompt Injection Plan and Schedule a Full Week of Threads Content From One Claude Conversation Coding Cat Oran Ep3, Five Tables Changed Everything Updated: BFF Pattern I'm done watching freelancers get buried by 200 proposals. So I'm building the alternative. This is my first post BFS Algorithm in Java Step by Step Tutorial with Examples Tracking LLM Pricing Monthly: An Open Dataset for 22 AI Models How We Measure Content ROI on a Comparison Site: Revenue Attribution Without Perfect Data Introducing Nova AI Ops: The AI-Native Operating System for SRE Teams I built a free desktop video downloader for Windows — Grabbit How Talkie OCR Helps Vision-Impaired & Dyslexic Users Read the World Around Them VRCFaceTracking安装和iPhone面捕配置教程,有bug Even CrowdStrike Can't See Your Agents The Automation Gold Rush: What n8n Workflows and Claude Are Opening Up for Developers Right Now
Your pipeline isn't slow. Your parser is.
Vinicius Fagundes · 2026-06-18 · via DEV Community

8 hours. That's how long a nightly batch job took when I picked it up two weeks ago. The team had already thrown a bigger cluster at it. Twice. It barely moved.

By the end it ran in 47 minutes. Same cluster. Same algorithms. I didn't touch the "smart" part of the pipeline at all.

The whole win was in how the data entered the system. The parser. The first few stages nobody ever profiles because they look boring.

This post is the long version of that fix — written so that if you've only been doing this for a year, you can follow every line. We'll define the jargon as we go, measure before we touch anything, and then walk through the five things that are almost certainly eating your runtime, each with before-and-after code you can run.

First, what is "the layer"? (30 seconds for newcomers)

A data pipeline is just a chain of stages. Raw data goes in one end, something useful comes out the other:

raw files  ->  parse  ->  transform  ->  aggregate / model  ->  output

  • Parse (also called ingestion or deserialization) means turning raw text — a JSON file, a CSV, a log line — into objects your program can work with, like Python dictionaries and lists. The bytes {"id": 7} on disk become the dict {"id": 7} in memory. That conversion is not free.
  • Transform is where you clean, reshape, join, and compute features.
  • Model / aggregate is the part people think of as "the real work" — the algorithm, the SQL, the math.

Here's the thing almost nobody internalizes early: that first box, parse, often costs more than all the others combined. And because it sits at the front, everything downstream pays its tax. A slow parser doesn't just slow down parsing — it slows down every stage that waits on it.

So when a pipeline is slow, the instinct is to make the smart parts faster. Usually the smart parts were never the problem.

Rule 1: never optimize what you haven't measured

Before you change a single line, find out where the time actually goes. This is the most important habit in this whole post, and it's three lines of code.

import time
from contextlib import contextmanager

@contextmanager
def stage(name):
    start = time.perf_counter()
    yield
    print(f"{name:<18} {time.perf_counter() - start:6.2f}s")

Now wrap each stage of your pipeline in it:

import json

with stage("read + parse"):
    records = [json.loads(line) for line in open("events.jsonl")]

with stage("transform"):
    rows = [transform(r) for r in records]

with stage("aggregate"):
    result = aggregate(rows)

You'll see something like this, and it's the whole story:

read + parse        372.10s
transform            48.30s
aggregate            12.05s

Parsing is 86% of the runtime. The "data science" everyone wanted to optimize is the bottom two lines.

When you need more detail — which function inside parsing is slow — reach for the profiler that ships with Python:

import cProfile
cProfile.run("run_pipeline()", sort="cumulative")

Read the top few rows of the output. If you see names like loads, decode, raw_decode, or scanstring near the top, that's the JSON parser, and you've just confirmed your bottleneck is ingestion — not the cluster, not Spark, not your SQL.

Everything below is what to do once the profiler points at parsing.

The five things quietly eating your night

1. You parse the same data five times

This is the most common one, and it hides because each stage looks reasonable on its own. The bug only shows up when you zoom out.

def clean_stage():
    data = [json.loads(l) for l in open("events.jsonl")]   # parse #1
    return [r for r in data if r["status"] == "ok"]

def feature_stage():
    data = [json.loads(l) for l in open("events.jsonl")]   # parse #2 (!)
    return build_features(data)

def report_stage():
    data = [json.loads(l) for l in open("events.jsonl")]   # parse #3 (!!)
    return summarize(data)

Three functions, each independently re-reading and re-parsing the same file from scratch. If parsing is 80% of your cost, you just paid it three times. I've seen pipelines pay it seven or eight times across a DAG because each task was written by a different person on a different day.

The fix is embarrassingly simple: parse once, pass the result down.

records = [json.loads(l) for l in open("events.jsonl")]   # parse ONCE

cleaned   = clean_stage(records)
features  = feature_stage(cleaned)
report    = report_stage(features)

In a real orchestrator (Airflow, dbt, Spark), "parse once" usually means land the parsed data somewhere — a temp table, a cached DataFrame, a Parquet file — so the next task reads the cheap version instead of re-parsing the raw source. More on that in #5.

2. You hydrate the whole object to use three fields

Say each record is a big nested blob, but your job only ever touches three values.

def parse(line):
    obj = json.loads(line)   # builds the ENTIRE nested tree in memory
    return obj               # ...and you carry all of it downstream

You can't avoid parsing the line to find your fields — but you can avoid carrying the whole thing. Pull out what you use and drop the rest at the boundary:

import orjson  # a C-based JSON parser, typically several times faster than the stdlib

def parse(line):
    obj = orjson.loads(line)
    return {
        "user_id": obj["user"]["id"],
        "ts":      obj["event"]["timestamp"],
        "amount":  obj["event"]["payload"]["amount"],
    }

Two wins here, and it's worth being precise about which is which:

  • orjson swaps the slow pure-Python decode path for a fast C one. That's a real parse-time speedup, and it's a one-line change — pip install orjson, change json to orjson.
  • The field projection doesn't make the parse cheaper (you still parse the line). It makes everything after the parse cheaper: less memory held, smaller objects to shuffle and serialize between stages, and no repeated deep lookups like obj["event"]["payload"]["amount"] happening over and over in later code. You do the deep walk once, here, and hand downstream a flat little dict.

Small objects move faster through every stage. A flat 3-key dict is cheaper to pass to Spark, pickle to a worker, or write to disk than a 200-key nested tree.

3. You load a 10 GB file just to read it once

If a single JSON document is bigger than your RAM, json.loads(f.read()) will either crash or swap to disk and crawl. You don't need the whole thing in memory at once — you need one record at a time. That's called streaming, and ijson does it:

import ijson

with open("huge.json", "rb") as f:
    for record in ijson.items(f, "item"):   # yields one array element at a time
        process(record)

ijson walks the file as a stream and hands you records as it finds them, so memory stays flat whether the file is 100 MB or 100 GB. The "item" argument is the path to the elements you want — for a top-level JSON array [ {...}, {...} ], each element is item.

If your data is newline-delimited JSON (one object per line — the .jsonl format), you don't even need a library; just iterate the file line by line and parse each line, which is already streaming:

with open("events.jsonl") as f:
    for line in f:            # one line in memory at a time
        record = orjson.loads(line)
        process(record)

The lesson: never pull a whole dataset into memory to look at it piece by piece.

4. You let Spark guess your schema

This one is a Spark-specific tax that surprises a lot of people. When you do this:

df = spark.read.json("s3://bucket/events/")

Spark doesn't know the structure of your JSON, so before it can read the data it launches an extra pass over the files just to figure out the schema — what columns exist and what types they are. Then it reads everything again to actually load it. On a big dataset that inference pass is a full, expensive scan you're paying for silently, and on wide/nested JSON it can also blow up driver memory.

Tell Spark the schema up front and it skips the guessing pass entirely:

from pyspark.sql.types import StructType, StructField, StringType, LongType, DoubleType

schema = StructType([
    StructField("user_id", StringType()),
    StructField("ts",      LongType()),
    StructField("amount",  DoubleType()),
])

df = spark.read.schema(schema).json("s3://bucket/events/")

Same result, one less scan, lower memory, and a bonus: an explicit schema means a malformed record fails loudly instead of silently inferring the wrong type and poisoning everything downstream. (The same idea applies to spark.read.csv(..., inferSchema=True) — inference always costs an extra read.)

5. The real fix: JSON is the wrong shape to begin with

Everything above is treating symptoms. Here's the cause.

JSON is row-oriented and untyped text. Every time anything reads it, it has to be parsed again from scratch, and the types re-derived. There is no way to read "just the amount column" — to get one field out of every record, the parser still has to chew through every record. You pay the full parse cost on every single read, forever.

The fix is to pay that cost exactly once, at ingestion, and convert the raw JSON into a columnar, typed format. The standard choice is Parquet.

import pyarrow.json as paj
import pyarrow.parquet as pq

# At ingestion, ONCE: parse the JSON a single time and write Parquet
table = paj.read_json("events.jsonl")
pq.write_table(table, "events.parquet")

From then on, every downstream job reads the Parquet, and reading Parquet is a different universe:

import pyarrow.parquet as pq

# reads ONLY these two columns off disk; the rest is never touched
df = pq.read_table("events.parquet", columns=["user_id", "amount"])

Why this is so much faster, in plain terms:

  • Columnar means values for each column are stored together, instead of record-by-record. So "give me the amount column" reads one contiguous chunk and ignores the other 197 columns. In row-oriented JSON, the same request has to read and parse every record in full.
  • Typed means the column already knows it's a 64-bit integer. No text-to-number parsing on read. The types are stored in the file.
  • Compressed + indexed means Parquet keeps min/max stats per chunk ("row group"), so a filter like amount > 1000 can skip entire chunks without reading them. That's called predicate pushdown — the filter is pushed down into the read itself.

The mental model: JSON makes every reader redo the work. Parquet does the work once and lets every reader skip to exactly what it needs. If your raw data lands as JSON and gets read more than once — and it always does — converting it to columnar at the front door is usually the single highest-leverage change you can make.

The pattern behind all five: normalize once at the front door

Step back and every fix above is the same move. Do the expensive, messy work — parsing, type-casting, field extraction, schema enforcement — once, at the boundary where data enters your system. Produce one clean, typed, columnar version. Then let every downstream stage read that cheap version instead of fighting the raw input over and over.

raw JSON  ->  [ parse + flatten + type + write Parquet ]  ->  clean columnar table
                          (pay this ONCE)                          |
                                                                   v
                                          every model / report / job reads this

This is why the 8-hour job became a 47-minute job without touching the algorithms. The old version re-parsed raw nested JSON inside every stage, inferred the schema on every read, and dragged the full objects through the whole DAG. The new version parsed once, landed typed Parquet, and let everything downstream read flat typed columns. Nothing got smarter. The work just moved to the layer it belonged in — and stopped being repeated.

The principle

If your parser is wrong, every layer downstream pays for it. Your model is only as fast as the data feeding it, and your data science is only as good as the shape the data arrives in. You can have a perfect algorithm sitting behind a parser that re-reads 10 GB of JSON five times a night, and you will spend your week buying bigger clusters to fix a problem that lives in the first ten lines of the job.

So before you reach for more compute, profile the thing and find out where the time actually goes. Nine times out of ten it's hiding at the front door, in the boring stage nobody looks at.

Last time you chased a slow pipeline — when you finally measured it, where did the time actually turn out to be hiding? I collect these stories.


I'm Vinicius Fagundes — principal data engineer, independent, and an MBA lecturer in São Paulo. I spend my days making slow, expensive, fragile pipelines fast, cheap, and boring. If this sounds like your stack, this is the work I do at vf-insights.com.