惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

C
CXSECURITY Database RSS Feed - CXSecurity.com
Security Latest
Security Latest
V
Vulnerabilities – Threatpost
Exploit-DB.com RSS Feed
Exploit-DB.com RSS Feed
L
LINUX DO - 最新话题
N
News and Events Feed by Topic
P
Proofpoint News Feed
G
GRAHAM CLULEY
NISL@THU
NISL@THU
S
Securelist
Recent Commits to openclaw:main
Recent Commits to openclaw:main
Schneier on Security
Schneier on Security
www.infosecurity-magazine.com
www.infosecurity-magazine.com
Application and Cybersecurity Blog
Application and Cybersecurity Blog
T
Threat Research - Cisco Blogs
Threat Intelligence Blog | Flashpoint
Threat Intelligence Blog | Flashpoint
T
Tenable Blog
Google DeepMind News
Google DeepMind News
Hacker News - Newest:
Hacker News - Newest: "LLM"
Help Net Security
Help Net Security
C
CERT Recently Published Vulnerability Notes
T
The Exploit Database - CXSecurity.com
I
Intezer
阮一峰的网络日志
阮一峰的网络日志
AI
AI
PCI Perspectives
PCI Perspectives
Hacker News: Ask HN
Hacker News: Ask HN
aimingoo的专栏
aimingoo的专栏
宝玉的分享
宝玉的分享
博客园 - 【当耐特】
腾讯CDC
The Cloudflare Blog
Spread Privacy
Spread Privacy
Latest news
Latest news
有赞技术团队
有赞技术团队
D
Docker
Cyberwarzone
Cyberwarzone
L
LINUX DO - 热门话题
S
Secure Thoughts
Hugging Face - Blog
Hugging Face - Blog
K
Kaspersky official blog
M
MIT News - Artificial intelligence
N
News | PayPal Newsroom
H
Hackread – Cybersecurity News, Data Breaches, AI and More
S
SegmentFault 最新的问题
J
Java Code Geeks
S
Security Affairs
The Register - Security
The Register - Security
云风的 BLOG
云风的 BLOG
Cyber Security Advisories - MS-ISAC
Cyber Security Advisories - MS-ISAC

DEV Community

Authentication Security Deep Dive: From Brute Force to Salted Hashing (With Java Examples) Why AI Systems Don’t Fail — They Drift Spilling beans for how i learn for exam😁"Reinforcement Learning Cheat Sheet" I Replaced Chrome with Safari for AI Browser Automation. Here's What Broke (and What Finally Worked) How Python Borrows Other People's Work The $40 Architecture: Processing 1 Billion API Requests with 99.99% Uptime Vibe Coding: A Workflow Guide (From Zero to SaaS) Most webhook security guides protect the wrong side. The scary part is delivery. Headless CMS for TanStack Start: Build a Blog with Cosmic EU Age Verification App "Hacked in 2 Minutes" — What Actually Happened Comfy Cloud’s delete function does not actually remove files Running AI Models on GPU Cloud Servers: A Beginner Guide Event-driven media intelligence with AWS Step Functions and Bedrock I scored 500 AI prompts across 8 quality dimensions — here's what broke How to Call Google Gemini API from Next.js (Free Tier, No Backend Needed) The Portal Protocol: Reclaiming Human Connection in the Age of AI How to Fix Your Team's Scattered Knowledge Problem With a Self-Hosted Forum Intro to tc Cloud Functors: A Graph-First Mental Model for the Modern Cloud Designing Multi-Tenant Backends With Both Ownership and Team Access I Built a Neumorphic CSS Library with 77+ Components — Here's What I Learned PostgreSQL Performance Optimization: Why Connection Pooling Is Critical at Scale Cómo construí un SaaS multi-rubro para gestionar expensas en Argentina con FastAPI + Vue 3 🚀 I Built an Ethical Hacking Scanner Tool – Open Source Project I Replaced /usage and /context in Claude Code With a Single Statusline A Pythonic Way to Handle Emails (IMAP/SMTP) with Auto-Discovery and AI-Ready Design I Collected 8.9 Million Polymarket Price Points — Here's What I Found About How Markets Really Move EcoTrack AI — Carbon Footprint Tracker & Dashboard Everyone's Using AI. No One Agrees How. 5 self-hosted ebook managers worth trying in 2026 Building Your First AI Agent with LangChain: From Chatbot to Autonomous Assistant Common SOC 2 Failures (Real World) Stop Vibe-Checking Your AI App: A Practical Guide to Evals How to Use SonarQube and SonarScanner Locally to Level Up Your Code Quality Your Next To-Do App Is Dead — I Replaced Mine with an OpenClaw AI Sign a Nostr event in 60 lines of Python using coincurve — no nostr-sdk, no nbxplorer, no rust toolchain ITGC Audit Explained Like You’re in Big 4 Patch Tuesday abril 2026: Microsoft parcha 163 vulnerabilidades y un zero-day en SharePoint Stop scraping everything: a better way to track competitor price changes Listing on MCPize + the Official MCP Registry while routing payments OUTSIDE the marketplace — how I kept 100% of my x402 revenue Building an AI-Powered Risk Intelligence System Using Serverless Architecture Why We Ripped Function Overloading Out of Our AI Toolchain Testing AI-Generated Code: How to Actually Know If It Works SaaS Churn Is Killing Your Business. Here Is What to Do About It (Without a Support Team) The Speed of AI Is No Longer Linear - And Self-Improving Models Are Why How to Implement RBAC for MCP Tools: A Practical Guide for Engineering Teams From Standard Quote to Persuasive Proposal: AI Automation for Arborists I built a CLI that scaffolds complete multi-tenant SaaS apps Axios CVE-2025–62718: The Silent SSRF Bug That Could Be Hiding in Your Node.js App Right Now The dashboard that ended our friendship Data Pipelines Explained Simply (and How to Build Them with Python) The Hidden Cost of AI Systems Nobody Talks About. undefined vs undeclared, and how typeof behaves Switching from file-based jobs to NATS/Kafka in Rust without changing code io_uring Adventures: Rust Servers That Love Syscalls Why Agentic AI is Killing the Traditional Database The POUR principles of web accessibility for developers and designers Quantum Neural Network 3D — A Deep Dive into Interactive WebGL Visualization How To Install Caveman In Codex On macOS And Windows Automation Pipeline Reliability: Why Your Workflow Breaks When Nobody Is Watching I Built an 'Open World' AI Coding Agent — It Works From ANY Folder From Freelancing to Product: A Tech Service Company's SaaS Transformation China's AI Giants: Adding Tencent Hunyuan & ByteDance Doubao to AI University (74 Providers) On the Vibe Coders and Their Lies clerk: Auto-Summarize Your Claude Code Sessions AI Weekly — 2026/04/10–04/17 | The Model Lockdown Is Here, but the Toolchain Is the Real Battleground AI 週報 — 2026/04/10–2026/04/17 模型封鎖潮來了,但工具鏈才是真戰場 Maybe this is how Open-Source apps are born... 🚀 Fine-Tune LLMs with LoRA and QLoRA: 2026 Guide tRPC v11 + Next.js App Router: End-to-End Type Safety Without the Boilerplate ShadCN UI in 2026: Why I Stopped Installing Component Libraries and Started Owning My Components SaaS Billing in React Server Components: Stripe + Supabase Without a Single `useEffect` Join our DEV Weekend Challenge — $1,000 in Prizes Across TEN winners! Submissions Due April 20 at 6:59 AM UTC. Implementing FSRS Spaced Repetition in Flutter + Supabase — Adding Memory Science to an AI Learning App "I Texted My Localhost From the Train — Claude Code Fixed the Bug Before I Got Home" I Built a Sales Prep AI and It Went Deeper Than Expected Design to Code #2: One JSON, Eleven Outputs Solving the 100M-Row Problem: A Summary Table Pattern for High-Volume Push Notification Logs Flutter Web With Wasm: What Actually Changes For Developers I Built 50 Royalty-Free Soundtracks for My Side Project in a Weekend Using AI Music Generation The Vibe Coding Security Checklist: 7 Things to Check Before You Ship Stop Letting Googlebot Guess Fix Your React App's SEO Right Desconstruindo o Streaming do LinkedIn: Como Criar um Engine de Extração de Vídeo de Alta Performance com HLS e FFmpeg (EDA Part-1) EDA (Exploratory Data Analysis) Explained With Real Life — Why Looking at Your Data Is the Most Important Step in Machine Learning Brand Relationship Management at Scale: Our 4-Touch Outreach System for 200+ Brands Why String.fromEnvironment() Might Return an Empty String in Dart JGuardrails 1.0.0 — Hardening Java LLM Apps Against Jailbreaks, Toxicity, and Prompt Injection Plan and Schedule a Full Week of Threads Content From One Claude Conversation Coding Cat Oran Ep3, Five Tables Changed Everything Updated: BFF Pattern I'm done watching freelancers get buried by 200 proposals. So I'm building the alternative. This is my first post BFS Algorithm in Java Step by Step Tutorial with Examples Tracking LLM Pricing Monthly: An Open Dataset for 22 AI Models How We Measure Content ROI on a Comparison Site: Revenue Attribution Without Perfect Data Introducing Nova AI Ops: The AI-Native Operating System for SRE Teams I built a free desktop video downloader for Windows — Grabbit How Talkie OCR Helps Vision-Impaired & Dyslexic Users Read the World Around Them VRCFaceTracking安装和iPhone面捕配置教程,有bug Even CrowdStrike Can't See Your Agents The Automation Gold Rush: What n8n Workflows and Claude Are Opening Up for Developers Right Now
Dependency Gap in AI Code: Declared 1, Imported 4
Alexey Spinov · 2026-06-24 · via DEV Community

The dependency gap in AI-generated code is what the source imports minus what the manifest declares. A green CI run only proves the packages were already installed on the author's machine, not on a fresh checkout. So measure the gap statically, before merge. repro_probe.py reads the source with ast and never runs it. On a project that declared one package but imported four: gap=3, coverage 25.0%, exit 1.

AI disclosure: I drafted this with an AI writing assistant. The tool, the three fixtures, and every number below come from a real local run on Python 3.13.5, stdlib only, no network. I ran it, checked the exit codes, hashed the STDOUT twice to confirm it's byte-for-byte deterministic, and edited every line before publishing.

Green CI is a machine telling you it agreed with itself. The interesting question is what happens on a machine that isn't yours.

Here's the failure I keep seeing. An agent writes a feature. It imports pandas, numpy, a YAML parser. The tests pass, because the agent's environment already had those installed from some earlier task. The diff lands. A teammate pulls it, runs pip install -r requirements.txt, runs the code, and gets ModuleNotFoundError: No module named 'numpy'. The manifest never learned about the import. Nobody lied. The information just never made it from the source file into the dependency list.

That gap is small, boring, and statically measurable. So I measured it.

What I actually ran

TL;DR.

  • A green checkmark proves the suite ran where the packages were already present. It does not prove the project installs from its own manifest.
  • repro_probe.py walks every *.py with the stdlib ast, reads the manifest as text, and never imports, installs, or runs anything in the target project.
  • Four deterministic rules per import: R1 stdlib, R2 local module, R3 declared-match, R4 undeclared third-party (the gap).
  • On a project that declared requests and imported four third-party packages: gap=3 (numpy, pandas, yaml), coverage 25.0%, exit 1.
  • On an honest project where every import is declared, local, or stdlib: gap=0, coverage 100.0%, exit 0. The claim is falsifiable and it passed the honest case.
  • STDOUT is byte-for-byte identical across two runs (sha256 matched). No key, no network. No manifest → exit 2.

The run comes before the argument, because the run is the argument.

The contrarian bit: passing is not reproducible

Most discussion of AI-generated code stops at "does it run." Did the agent produce something? Did the tests go green? Ship it. The expensive part is the part nobody checks at merge time: would this install and import on a clean machine, from nothing but its own manifest?

There's recent evidence the gap is real and not rare. In AI-Generated Code Is Not Reproducible (Yet) (arXiv 2512.22387, v3, March 2026), Vangala, Adibifar, Gehani and Malik generated 300 projects from 100 prompts across Claude Code, OpenAI Codex and Gemini, then tried to run each one. Their abstract reports that only 68.3% of projects execute out-of-the-box, so roughly a third fail on first run. By language the spread is wide: Python 89.2%, Java 44.0%. And the line that made me build this tool: they measured a 13.5x average expansion from declared to actual runtime dependencies. Declared three, needed dozens. That is exactly the shape my broken fixture imitates, just smaller.

I want to be careful with those numbers. They are their measurement, not mine: a controlled study of generated projects, cited here with the source so you can read the methodology yourself. My own numbers below are only from my fixtures.

The four rules

The whole thing is one file. Every import in the source gets sorted into exactly one bucket, deterministically:

  • R1 stdlib. The import is in sys.stdlib_module_names (Python 3.10+ ships this set). os, json, pathlib: no declaration needed.
  • R2 local. There's a name.py or name/__init__.py in the project. It's your own module, not a dependency.
  • R3 declared-match. The normalized import name is in the manifest. This is where the annoying cases live: you import yaml but the package is PyYAML; you import cv2 but install opencv-python. A small map handles the common ones.
  • R4 undeclared. Not stdlib, not local, not declared. That's the gap. It imports a third-party package the manifest never mentions, so a fresh pip install won't pull it.

The metric is just the size of the R4 set. Here's the classification loop, verbatim:

for name in sorted(imports):
    if name in std:
        rule = "R1 stdlib"
    elif name in local:
        rule = "R2 local"
    elif norm(DIST_MAP.get(name, name)) in decl:
        rule = "R3 declared"
    else:
        rule = "R4 UNDECLARED"; gap.append(name)
    verdicts.append((name, rule))

No model call. No pip. No subprocess. It reads text and walks an AST.

The broken project: declared one, imported four

The fixture is a tiny "agent-style" file. It imports requests, numpy, pandas, and yaml, plus a local utils module and two stdlib modules. The requirements.txt has exactly one line: requests. Here is the real, unedited output:

repro_probe v1 | project=broken_project
  logging          R1 stdlib
  numpy            R4 UNDECLARED
  os               R1 stdlib
  pandas           R4 UNDECLARED
  requests         R3 declared
  utils            R2 local
  yaml             R4 UNDECLARED
reproducibility_gap=3 declared_coverage=25.0% gate=0
imported_but_not_declared=numpy,pandas,yaml
verdict=BROKEN exit=1

Three undeclared third-party imports out of four. Coverage 25.0%. Exit 1, which in CI fails the build.

Look at yaml. It's flagged R4 here because the manifest doesn't list its distribution name PyYAML. That's the import-name vs distribution-name trap that makes grep import so unreliable: the strings don't match even when the dependency is "obvious." The tool normalizes through its map, so it catches the case a naive text search would miss, and (as the next section shows) clears it when PyYAML actually is declared.

The honest project: the claim has to be falsifiable

A check that flags everything is worthless. If my contrarian line ("passing is not reproducible") is real, then a genuinely clean project must come back clean. So the second fixture declares every third-party import it uses, imports a local helpers module, and leans on stdlib. Same tool, same rules:

repro_probe v1 | project=clean_project
  helpers          R2 local
  io               R1 stdlib
  json             R1 stdlib
  os               R1 stdlib
  pathlib          R1 stdlib
  requests         R3 declared
  yaml             R3 declared
reproducibility_gap=0 declared_coverage=100.0% gate=0
imported_but_not_declared=(none)
verdict=CLEAN exit=0

Gap 0, coverage 100.0%, exit 0. Note yaml is R3 declared here. Same import, opposite verdict, because this manifest lists PyYAML. The map cuts both ways: it doesn't false-flag a dependency that's declared under its real distribution name. The honest project passes. That's the part that makes the tool a check and not a rubber stamp.

The third fixture has source but no requirements.txt and no pyproject.toml. It exits 2. Bad input, refuse to guess. A tool that invented a verdict from a missing manifest would be worse than no tool.

Is it deterministic?

A pre-merge gate that flickers is noise. I ran each fixture twice and hashed STDOUT only (service messages go to stderr, kept out of the hashed stream):

clean_project   run1=5177ac0a...  run2=5177ac0a...  -> IDENTICAL
broken_project  run1=52a677f9...  run2=52a677f9...  -> IDENTICAL
bad_project     run1=e3b0c442...  run2=e3b0c442...  -> IDENTICAL

Same input, same bytes, every time. The output is sorted, so there's no set-ordering wobble. You can wire the exit code into CI and trust that a clean tree stays green.

What this is NOT

This is the part that keeps the tool honest, so read it before you quote a number.

A dependency gap is a reproducibility signal, not proof the project won't run. It says "a third-party package is imported but not declared," which is a strong reason a fresh install would fail, but not a guarantee. Maybe the package is provided by the base image. Maybe it's an optional code path. The gap flags risk; it doesn't render a verdict on the project's fate. I called the metric gap, not brokenness, on purpose.

And it has real blind spots I'm not going to hide:

  • It doesn't check version pins. numpy declared but pinned to a version that doesn't exist, or a range that conflicts, sails through as R3. The gap is about presence, not resolvability.
  • It doesn't follow transitive dependencies. It sees your direct imports, not what those packages drag in. The 13.5x expansion the paper measured lives mostly in the transitive layer, which a static import-vs-manifest check can't reach.
  • It doesn't understand extras or environment markers. package[extra] and ; sys_platform == ... are normalized down to the base name; the extra is not verified.
  • The stdlib allowlist is version-bound. It's whatever sys.stdlib_module_names reports on the Python you run it with. Run it on a different minor version and a module could shift buckets.
  • It only reads requirements.txt and PEP 621 [project] dependencies. Poetry's [tool.poetry.dependencies], setup.py/setup.cfg, and optional-dependencies are not parsed as the dependency set. Point it at a poetry project and every import comes back undeclared, a false BROKEN, not a real gap. Run it on requirements-based projects, or read the verdict knowing this.
  • Local detection is top-level only. It treats name.py or name/__init__.py in the project root as local (R2). A src/-layout package or a PEP 420 namespace package (a local dir with no __init__.py) gets flagged R4 instead. That's a false positive on a perfectly reproducible project, not a missing dependency.

These false-BROKEN cases all err the safe way for a gate (it complains too much, never too little), but they're why you read the named list, not just the exit code, before you trust a verdict.

So: a clean exit 0 is necessary, not sufficient. It rules out the dumbest, most common failure (imported but never declared) and nothing more. That one class is worth ruling out, because it's the one that turns a green PR into a teammate's ModuleNotFoundError.

How this sits next to the other checks

This isn't a runtime guard. It's an artifact check that runs before the merge button, on the diff, with nothing executing. It's a cousin of the green-checkmark auditor, which asks whether passing tests carry independent signal; here the question is whether the manifest matches the imports. Both come from the same suspicion that a green status is a claim, not a proof. It's the same suspicion behind your agent returns 200 and lies. If you already parse manifests for drift, it rhymes with pinning and verifying MCP tool manifests. And the whole instinct, put the barrier before the irreversible step rather than after, is the pre-execution gate applied to the merge instead of the runtime.

Run it on your own repo

The tool is one Python file, stdlib only. Point it at a real project:

python3 repro_probe.py path/to/your/project
echo "exit: $?"

Exit 0 means every imported third-party package is declared (or it's stdlib/local). Exit 1 hands you the list of imported-but-not-declared names. Exit 2 means there's no manifest to check against. Add --gate N if you want to tolerate a known gap of N while you fix it.

It won't catch a bad version pin or a missing transitive dep. It will catch the most common reason an AI-written diff passes CI and then won't install: the package the source imports and the manifest forgot.

If you run it on something an agent wrote recently, I'd genuinely like to know the gap you got, and whether any of it was the import-name vs distribution-name trap. That's the case I'm least sure my little map covers well. Follow along for the next batch of numbers, and drop your worst gap in the comments.