惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

Hacker News - Newest:
Hacker News - Newest: "LLM"
The Last Watchdog
The Last Watchdog
L
LINUX DO - 最新话题
Application and Cybersecurity Blog
Application and Cybersecurity Blog
T
Troy Hunt's Blog
Cloudbric
Cloudbric
N
News | PayPal Newsroom
Security Archives - TechRepublic
Security Archives - TechRepublic
TaoSecurity Blog
TaoSecurity Blog
H
Hacker News: Front Page
Help Net Security
Help Net Security
S
Secure Thoughts
Recent Commits to openclaw:main
Recent Commits to openclaw:main
Exploit-DB.com RSS Feed
Exploit-DB.com RSS Feed
cs.CV updates on arXiv.org
cs.CV updates on arXiv.org
PCI Perspectives
PCI Perspectives
AI
AI
Hacker News: Ask HN
Hacker News: Ask HN
NISL@THU
NISL@THU
Last Week in AI
Last Week in AI
Forbes - Security
Forbes - Security
The GitHub Blog
The GitHub Blog
D
DataBreaches.Net
Scott Helme
Scott Helme
Jina AI
Jina AI
T
Threatpost
W
WeLiveSecurity
P
Palo Alto Networks Blog
F
Fortinet All Blogs
腾讯CDC
人人都是产品经理
人人都是产品经理
云风的 BLOG
云风的 BLOG
博客园 - 【当耐特】
Apple Machine Learning Research
Apple Machine Learning Research
P
Privacy International News Feed
P
Proofpoint News Feed
K
KPMG report finds enterprise disconnect between AI and its ROI | CIO
A
About on SuperTechFans
V
Vulnerabilities – Threatpost
www.infosecurity-magazine.com
www.infosecurity-magazine.com
C
Cyber Attacks, Cyber Crime and Cyber Security
B
Blog
Cyber Security Advisories - MS-ISAC
Cyber Security Advisories - MS-ISAC
D
Darknet – Hacking Tools, Hacker News & Cyber Security
IT之家
IT之家
美团技术团队
I
InfoQ
阮一峰的网络日志
阮一峰的网络日志
T
Threat Research - Cisco Blogs
博客园 - 司徒正美

DEV Community

Authentication Security Deep Dive: From Brute Force to Salted Hashing (With Java Examples) Why AI Systems Don’t Fail — They Drift Spilling beans for how i learn for exam😁"Reinforcement Learning Cheat Sheet" I Replaced Chrome with Safari for AI Browser Automation. Here's What Broke (and What Finally Worked) How Python Borrows Other People's Work The $40 Architecture: Processing 1 Billion API Requests with 99.99% Uptime Vibe Coding: A Workflow Guide (From Zero to SaaS) Most webhook security guides protect the wrong side. The scary part is delivery. Headless CMS for TanStack Start: Build a Blog with Cosmic EU Age Verification App "Hacked in 2 Minutes" — What Actually Happened Comfy Cloud’s delete function does not actually remove files Running AI Models on GPU Cloud Servers: A Beginner Guide Event-driven media intelligence with AWS Step Functions and Bedrock I scored 500 AI prompts across 8 quality dimensions — here's what broke How to Call Google Gemini API from Next.js (Free Tier, No Backend Needed) The Portal Protocol: Reclaiming Human Connection in the Age of AI How to Fix Your Team's Scattered Knowledge Problem With a Self-Hosted Forum Intro to tc Cloud Functors: A Graph-First Mental Model for the Modern Cloud Designing Multi-Tenant Backends With Both Ownership and Team Access I Built a Neumorphic CSS Library with 77+ Components — Here's What I Learned PostgreSQL Performance Optimization: Why Connection Pooling Is Critical at Scale Cómo construí un SaaS multi-rubro para gestionar expensas en Argentina con FastAPI + Vue 3 🚀 I Built an Ethical Hacking Scanner Tool – Open Source Project I Replaced /usage and /context in Claude Code With a Single Statusline A Pythonic Way to Handle Emails (IMAP/SMTP) with Auto-Discovery and AI-Ready Design I Collected 8.9 Million Polymarket Price Points — Here's What I Found About How Markets Really Move EcoTrack AI — Carbon Footprint Tracker & Dashboard Everyone's Using AI. No One Agrees How. 5 self-hosted ebook managers worth trying in 2026 Building Your First AI Agent with LangChain: From Chatbot to Autonomous Assistant Common SOC 2 Failures (Real World) Stop Vibe-Checking Your AI App: A Practical Guide to Evals How to Use SonarQube and SonarScanner Locally to Level Up Your Code Quality Your Next To-Do App Is Dead — I Replaced Mine with an OpenClaw AI Sign a Nostr event in 60 lines of Python using coincurve — no nostr-sdk, no nbxplorer, no rust toolchain ITGC Audit Explained Like You’re in Big 4 Patch Tuesday abril 2026: Microsoft parcha 163 vulnerabilidades y un zero-day en SharePoint Stop scraping everything: a better way to track competitor price changes Listing on MCPize + the Official MCP Registry while routing payments OUTSIDE the marketplace — how I kept 100% of my x402 revenue Building an AI-Powered Risk Intelligence System Using Serverless Architecture Why We Ripped Function Overloading Out of Our AI Toolchain Testing AI-Generated Code: How to Actually Know If It Works SaaS Churn Is Killing Your Business. Here Is What to Do About It (Without a Support Team) The Speed of AI Is No Longer Linear - And Self-Improving Models Are Why How to Implement RBAC for MCP Tools: A Practical Guide for Engineering Teams From Standard Quote to Persuasive Proposal: AI Automation for Arborists I built a CLI that scaffolds complete multi-tenant SaaS apps Axios CVE-2025–62718: The Silent SSRF Bug That Could Be Hiding in Your Node.js App Right Now The dashboard that ended our friendship Data Pipelines Explained Simply (and How to Build Them with Python) The Hidden Cost of AI Systems Nobody Talks About. undefined vs undeclared, and how typeof behaves Switching from file-based jobs to NATS/Kafka in Rust without changing code io_uring Adventures: Rust Servers That Love Syscalls Why Agentic AI is Killing the Traditional Database The POUR principles of web accessibility for developers and designers Quantum Neural Network 3D — A Deep Dive into Interactive WebGL Visualization How To Install Caveman In Codex On macOS And Windows Automation Pipeline Reliability: Why Your Workflow Breaks When Nobody Is Watching I Built an 'Open World' AI Coding Agent — It Works From ANY Folder From Freelancing to Product: A Tech Service Company's SaaS Transformation China's AI Giants: Adding Tencent Hunyuan & ByteDance Doubao to AI University (74 Providers) On the Vibe Coders and Their Lies clerk: Auto-Summarize Your Claude Code Sessions AI Weekly — 2026/04/10–04/17 | The Model Lockdown Is Here, but the Toolchain Is the Real Battleground AI 週報 — 2026/04/10–2026/04/17 模型封鎖潮來了,但工具鏈才是真戰場 Maybe this is how Open-Source apps are born... 🚀 Fine-Tune LLMs with LoRA and QLoRA: 2026 Guide tRPC v11 + Next.js App Router: End-to-End Type Safety Without the Boilerplate ShadCN UI in 2026: Why I Stopped Installing Component Libraries and Started Owning My Components SaaS Billing in React Server Components: Stripe + Supabase Without a Single `useEffect` Join our DEV Weekend Challenge — $1,000 in Prizes Across TEN winners! Submissions Due April 20 at 6:59 AM UTC. Implementing FSRS Spaced Repetition in Flutter + Supabase — Adding Memory Science to an AI Learning App "I Texted My Localhost From the Train — Claude Code Fixed the Bug Before I Got Home" I Built a Sales Prep AI and It Went Deeper Than Expected Design to Code #2: One JSON, Eleven Outputs Solving the 100M-Row Problem: A Summary Table Pattern for High-Volume Push Notification Logs Flutter Web With Wasm: What Actually Changes For Developers I Built 50 Royalty-Free Soundtracks for My Side Project in a Weekend Using AI Music Generation The Vibe Coding Security Checklist: 7 Things to Check Before You Ship Stop Letting Googlebot Guess Fix Your React App's SEO Right Desconstruindo o Streaming do LinkedIn: Como Criar um Engine de Extração de Vídeo de Alta Performance com HLS e FFmpeg (EDA Part-1) EDA (Exploratory Data Analysis) Explained With Real Life — Why Looking at Your Data Is the Most Important Step in Machine Learning Brand Relationship Management at Scale: Our 4-Touch Outreach System for 200+ Brands Why String.fromEnvironment() Might Return an Empty String in Dart JGuardrails 1.0.0 — Hardening Java LLM Apps Against Jailbreaks, Toxicity, and Prompt Injection Plan and Schedule a Full Week of Threads Content From One Claude Conversation Coding Cat Oran Ep3, Five Tables Changed Everything Updated: BFF Pattern I'm done watching freelancers get buried by 200 proposals. So I'm building the alternative. This is my first post BFS Algorithm in Java Step by Step Tutorial with Examples Tracking LLM Pricing Monthly: An Open Dataset for 22 AI Models How We Measure Content ROI on a Comparison Site: Revenue Attribution Without Perfect Data Introducing Nova AI Ops: The AI-Native Operating System for SRE Teams I built a free desktop video downloader for Windows — Grabbit How Talkie OCR Helps Vision-Impaired & Dyslexic Users Read the World Around Them VRCFaceTracking安装和iPhone面捕配置教程,有bug Even CrowdStrike Can't See Your Agents The Automation Gold Rush: What n8n Workflows and Claude Are Opening Up for Developers Right Now
The Five-Thousand-Line File
Ian Johnson · 2026-05-22 · via DEV Community

Every team has one. Sometimes it is called utils.ts or helpers.py. Sometimes it has the name of a domain concept that originally meant something specific and has since absorbed everything tangentially related. The file is large enough that nobody opens it casually. It has multiple maintainers, each of whom understands a different third of it. New additions go into it because that is where similar things already live, and the file gets larger.

This is the god file: a single file that has grown to do too much, and that resists the refactor that would split it because the refactor is large and the file mostly works.

The god file is one of the most agent-hostile shapes a codebase can take.

How files get this big

No file is born five thousand lines long. The growth is incremental.

The file starts as something reasonable: a module that handles one thing, three hundred lines, well-organized. A developer adds a function related to the thing. Another developer adds a function that is related to one of the existing functions, but also touches a new concept. The new concept does not justify its own file, so it lives in this one. Over time, the file accumulates concepts that share an author, a directory, or nothing in particular except convenience.

The reason nobody splits it is that splitting it is a project. The file is imported from many places. Each import has to be updated. The functions inside have shared private helpers that have to be sorted out. Tests for the file have the same problem in miniature. The estimate for the refactor is "a sprint," and a sprint is always too expensive when the file is "fine."

So the file stays. New developers add to it because the existing functions are there. The agent does the same. The file grows.

Why agents struggle here

The god file is expensive for an agent in three specific ways.

The first is context budget. The agent loads files into its working memory to understand them. A five-thousand-line file consumes a large fraction of that budget for a small change. The agent has less room left for the rest of the codebase — the calling files, the tests, the conventions. Quality drops, not because the agent is dumber, but because it is operating with less situational awareness.

The second is pattern dilution. The agent pattern-matches against the file it is editing. A file with five hundred coherent lines teaches the agent one strong pattern. A file with five thousand lines teaches the agent ten weak patterns, often contradictory. The agent picks one, often the wrong one for the specific change.

The third is the path-of-least-resistance problem. When asked to add new functionality, the agent looks for where similar functionality lives. It finds the god file. It adds to the god file. The file grows by one more function, in the same shape as the previous additions. The agent, like every previous contributor, has chosen the cheap path. The file is now slightly more god-like.

A small coherent file is a force multiplier for an agent. A god file is a tax.

When size is actually the problem

It is worth being careful about the diagnosis. Not every large file is a god file. A file that defines a complex but coherent thing (a state machine, a parser, a single algorithm) may legitimately be large. The size is not the smell. The smell is unrelated things sharing a file.

The diagnostic question is: if you had to give this file a name that described what it does, in a single concept, could you? parser.ts is a coherent file even at three thousand lines, because everything in it is parser. helpers.ts is incoherent at five hundred lines, because nothing about the name tells you what is in it. The size is downstream of the coherence.

A useful test: pick five functions from the file at random. Do they belong together? If yes, the file is big but legitimate. If no, the file is a junk drawer with a misleading name.

Splitting by concern

The right way to split a god file is not by line count. It is by concern.

Look at the functions in the file. Group them by what they are for, not by what they touch. Two functions that both manipulate strings are not necessarily related; two functions that both implement steps of the same workflow are.

For each group, ask: would this group, alone, make sense as a file? Does it have a name that describes what it does? Are the dependencies between this group and the rest of the file mostly external, or mostly internal?

Groups that score well on this become candidate files. Move them. The imports update mechanically. The tests follow. The original god file shrinks by one concept; the codebase gains a coherent module.

This is the kind of refactor an agent is good at, given a clear scope. "Move these eight functions to a new file called pricing.ts, update all callers, and split the corresponding test file." A concrete instruction. The agent does the mechanical work. A human reviews the result.

Limit the size mechanically

Once you have done the initial split, the way to keep the file from re-growing is the same as with every other limit: make it mechanical.

Most linters can enforce a maximum file length. Set the limit slightly above your current largest legitimate file. The build fails when a file exceeds it. New code cannot grow a file past the limit; it has to go somewhere else.

The limit is a forcing function, not a precise number. The point is not that 500 is correct and 501 is wrong. The point is that the team is forced to make an active decision when a file approaches the limit, instead of letting it drift past 1,000, 2,000, 5,000 without noticing.

The agent will respect the limit because the agent runs the linter. It will offer to put new functions in new files when the existing file is near the threshold. The default direction shifts from "grow the god file" to "split the god file," which is what you wanted.

First steps

If your codebase has god files and you want to start fixing them:

Find your largest source file. Count its lines. Note the number. Open it and look at the function list. Are they coherent? Or is the file a junk drawer?

If it is a junk drawer, pick the most distinct group of functions — the smallest set you can extract without untangling shared dependencies. Move that group to its own file. Update imports. Run tests. Ship the PR.

Add a max-lines rule to your linter, set 20% above the largest file you have decided to keep. The build now prevents new files from exceeding the limit.

Quarterly, look at the file-length distribution. Pick the largest file. Split one group. Repeat.

Add a rule to AGENTS.md: "When adding new functions, prefer creating a new file in a domain-appropriate directory over extending a large existing file. Files larger than [N] lines are a smell; do not extend them without splitting at the same time."

The god file did not arrive overnight. It will not leave overnight. But the trajectory matters. A team that splits one group per quarter is on a path toward a codebase made of coherent modules. A team that does not is on a path toward one file that contains everything, and an agent that gets worse the more it touches the codebase.

The size of any one file is small. The cost of letting them all grow is not.