惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

V
Visual Studio Blog
Exploit-DB.com RSS Feed
Exploit-DB.com RSS Feed
Hugging Face - Blog
Hugging Face - Blog
D
DataBreaches.Net
A
About on SuperTechFans
D
Docker
腾讯CDC
Google DeepMind News
Google DeepMind News
Hacker News - Newest:
Hacker News - Newest: "LLM"
W
WeLiveSecurity
Forbes - Security
Forbes - Security
S
Security @ Cisco Blogs
V
Vulnerabilities – Threatpost
cs.AI updates on arXiv.org
cs.AI updates on arXiv.org
Know Your Adversary
Know Your Adversary
钛媒体:引领未来商业与生活新知
钛媒体:引领未来商业与生活新知
I
InfoQ
P
Privacy International News Feed
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻
Application and Cybersecurity Blog
Application and Cybersecurity Blog
aimingoo的专栏
aimingoo的专栏
C
Cyber Attacks, Cyber Crime and Cyber Security
S
Secure Thoughts
Stack Overflow Blog
Stack Overflow Blog
T
Tenable Blog
T
Threatpost
P
Proofpoint News Feed
博客园 - 司徒正美
Microsoft Security Blog
Microsoft Security Blog
N
News and Events Feed by Topic
Apple Machine Learning Research
Apple Machine Learning Research
云风的 BLOG
云风的 BLOG
GbyAI
GbyAI
Hacker News: Ask HN
Hacker News: Ask HN
B
Blog
CTFtime.org: upcoming CTF events
CTFtime.org: upcoming CTF events
Microsoft Azure Blog
Microsoft Azure Blog
F
Fortinet All Blogs
Project Zero
Project Zero
Help Net Security
Help Net Security
WordPress大学
WordPress大学
F
Full Disclosure
博客园 - 三生石上(FineUI控件)
O
OpenAI News
cs.CV updates on arXiv.org
cs.CV updates on arXiv.org
H
Heimdal Security Blog
Security Latest
Security Latest
T
Tor Project blog
C
CXSECURITY Database RSS Feed - CXSecurity.com
Cisco Talos Blog
Cisco Talos Blog

Hacker News

Introducing Claude Opus 4.7 Qwen Studio The Future of Everything is Lies, I Guess: Where Do We Go From Here? GitHub - SeanFDZ/macmind: Single-layer transformer in HyperTalk for the classic Macintosh Show HN: Agent-cache – Multi-tier LLM/tool/session caching for Valkey and Redis Bonsai 1-bit WebGPU - a Hugging Face Space by webml-community Moving a large-scale metrics pipeline from StatsD to OpenTelemetry / Prometheus GitHub - Nightmare-Eclipse/RedSun: The Red Sun vulnerability repository GitHub - SethPyle376/hiraeth: Local AWS emulator focused on fast integration testing, with SQS support, SQLite-backed state, and a debug-friendly web UI. GitHub - macOS26/Agent: Any AI, replaces Claude Code, Cursor, OpenClaw. Over 18 LLM providers (Claude, OpenAI, Gemini, Ollama, Zai, HF, Qwen) wired into a native Mac app that writes code, builds Xcode projects, bumps versions, manages git, automates Safari, use AppleScript, JS or Accessibility, extend Agent! w/ MCP Servers, run tasks from your iPhone via Messages. YouTube now lets you turn off Shorts I Made a Terminal Pager Burgers | マクドナルド公式 Commands — HackerNews CLI documentation ChatGPT for Excel PiCore - Raspberry Pi Port of Tiny Core Linux Live Nation illegally monopolized ticketing market, jury finds Google Broke Its Promise to Me. Now ICE Has My Data. Founding Engineer at Adaptional | Y Combinator CRISPR takes important step toward silencing Down syndrome’s extra chromosome GitHub - saffron-health/libretto: The AI toolkit for building reliable browser automations US v. Heppner (S.D.N.Y. 2026) no attorney-client privilege for AI chats [pdf] Retrofitting JIT Compilers into C Interpreters IPv6 – Google The Accursèd Alphabetical Clock Cybersecurity Looks Like Proof of Work Now Fragments: April 14 Cal.com Goes Closed Source: Why AI Security Is Forcing Our Decision | Cal.com - Scheduling Software for Online Bookings Laravel raised money and now injects ads directly into your agent When moving fast, talking is the first thing to break Too much Discussion of the XOR swap trick – Heather Cafe Introduction to Spherical Harmonics for Graphics Programmers The Grand Line Building a Z-Machine in the worst possible language High-Level Rust: Getting 80% of the Benefits with 20% of the Pain GitHub - duguyue100/midnight-captain: Inspired by Midnight Commander, tailored to my taste. How to build a `git diff` driver · Jamie Tanna | Software Engineer Center for Responsible, Decentralized Intelligence at Berkeley The Local Universe’s Expansion Rate Is Clearer Than Ever, but Still Doesn’t Add Up - A new synthesis of astronomical measurements confirms a persistent mismatch that could point to physics beyond current models The air throughout our homes is infused with microplastics. But there are things you can do to breathe less of them The disturbing white paper Red Hat is trying to erase from the internet – OSnews The Future of Everything is Lies, I Guess: Annoyances ‘Abhorrent’: the inside story of the Polymarket gamblers betting millions on war Productive procrastination — Max van IJsselmuiden maps, territory and LMs 447 Terabytes per Square Centimetre at Zero Retention Energy: Non-Volatile Memory at the Atomic Scale on Fluorographane Show HN: Pardonned.com – A searchable database of US Pardons 20 Years on AWS and Never Not My Job The Seasons are Wrong Artemis II crew splashes down near San Diego after historic moon mission We gave an AI a 3 year retail lease in SF and asked it to make a profit | Andon Labs How a dancer with ALS used brainwaves to perform live On filing the corners off my MacBooks Installing every* Firefox extension OpenClaw’s memory is unreliable, and you don’t know when it will break Steve Blank Nowhere Is Safe Chimpanzees in Uganda locked in vicious 'civil war', say researchers watgo - a WebAssembly Toolkit for Go linux/Documentation/process/coding-assistants.rst at master · torvalds/linux GitHub - callumlocke/json-formatter: Makes JSON easy to read. Founding Product Engineer at Bild AI | Y Combinator A compelling title that is cryptic enough to get you to take action on it GitHub - Keychron/Keychron-Keyboards-Hardware-Design: Industrial design files for Keychron keyboards and mice. 100+ models with CAD assets in STEP, DXF, DWG, and PDF. Source-available, with commercial use allowed for original compatible accessories within the license terms. [ANNOUNCE] WireGuardNT v0.11 and WireGuard for Windows v0.6 Released 1D-Chess Helium Is Hard to Replace Cooperative Vectors Introduction | Evolve Keeping a Postgres queue healthy — PlanetScale Our response to the Axios developer tool compromise Do Americans read print books, e-books or audiobooks more? The Zettelkasten Method in Obsidian: A Practical Setup Guide Artemis II Is Competency Porn and We Are Starving For It WeakC4 Flight Viz — Cockpit View A Mexican surveillance giant you’ve never heard of is now watching the U.S. border Surelock: Deadlock-Free Mutexes for Rust RISC-V 101 – what is it and what does it mean for Canonical? | Ubuntu The Problem That Built an Industry How Much Linear Memory Access Is Enough? | Solidean Investigating Split Locks on x86-64 Simplest hash functions Sybilproof reputation mechanisms (2005) [pdf] What is a property? How Complex is my Code? Static code analysis in Kotlin — tools overview Toffoli gates are all you need PGLite evangelism dcmake: a new CMake debugger UI Clojure on Fennel part one: Persistent Data Structures Fragments: April 2 Python Release Python install manager 26.1 The Life and Death of the Book Review - Liberties Introducing Database Traffic Control — PlanetScale Bitcoin miners are losing $19,000 on every BTC produced as difficulty drops 7.8% God sleeps in the minerals Building slogbox Apple Silicon and Virtual Machines: Beating the 2 VM Limit Who was “Not Even Wrong” first? Pokemon Evolution Vs Darwinian Evolution The APL Programming Language Source Code
What I’m Finding About LLM Code Style and Token Costs
Jim Montgomery · 2026-06-25 · via Hacker News

Spending output tokens to share it. Before the price spikes.

Jim


Where This Started

I’ve been working through creating and reviewing features with Claude the past year. It’s been remarkable seeing the tension in token consumption and legacy patterns. Right when I think something is complete, a problem surfaces—regression, edge case, whatever. All the while watching the slow, steady and natural march toward eventual full-price rates. Alongside this phenomenon, my accumulated push to stay at the pragmatic edge of modern Web work. The sweet spot where nearly ubiquitous features remove lines of code and improve quality—the place where I keep wondering: why did I get that output? Why did that line of code appear instead of what’s been available for years? I usually dismiss it with the observable fact that Claude is effectively junior level at best, and a useful approximation of the encyclopedic knowledge asked in interviews.

In trying to make progress on something I am finding myself reviewing my practice and looking at where that outrageous token usage is coming from. Every one of those is output tokens, the ones that cost several times more (3x to 5x!!!) than input tokens in API pricing. Patterns that are longer, more fragile, more insecure, and solving problems the platform already solved–often years ago.

It’s enough to start imagining there’s some conspiracy to take the entire web platform backward, right when Ryan Dahl and separately Alex Russell, Dimitri Glazkov (and many others) made Web Components, etc. They literally made the entire Web platform great again. All to eke out some return on the tokens. So for the sake of conspiracy, this is what I’m finding.

Because my background as human being, who uses language, designed typography, programmed early on, alongside drawing and many other eclectic oddities, I actually consider things like tabs as a remarkable innovation. I can literally reduce indentation to 1 character, not some abstraction I have to go ask someone how to define or get permission to use. (I guess I’m just far too egalitarian to appreciate the exclusionary attitude of the entire software community.) I care about humans, and want things to work within some parsimonious baseline. And multiplying stuff by 4 or some arbitrary number just really doesn’t make sense–to me. I could go on, but maybe this grounds the orientation—someone who’s worked with actual language on actual media and has opinions about when something works and when it doesn’t. That part tends to speak for itself.

I mention this because it colors what I looked into from a purely pragmatic standpoint. I’m not arguing for a specific position where everyone uses tabs (despite that speaking for itself). I’m disclosing background that shaped opinions I’d been sitting on—there was always an economic argument I kept to myself, and it’s now showing up in real API costs. My opinions on convention are not the article. The token usage optimizations are what I came here to share. So you can benefit too. If you want to keep using multiple spaces, I’ll remind myself that the literature said it seemed ok and the LLM doesn’t know any better.

The Easiest Token Optimization on the Planet Is Already in the Runtime

Deno and runtimes like Cloudflare Workers implement the Web API surface nativelyURL, URLSearchParams, fetch, FormData, Headers, Request, Response, AbortController, ReadableStream, crypto, and more—the same objects that run in the browser. This is the architectural choice that Deno made deliberately, and that WinterCG has been formalizing as a minimum common API surface across runtimes and it has a significant practical consequence: the same API surface covers both browser and server-side code. No translation layer, no shims, no adaptation cost. The platform has already solved a large category of problems, correctly, securely, and without dependencies. Deno is particularly notable for including a standard library where something may be missing and needs cross-platform solutions.

The LLM doesn’t know this about your environment unless you say so. Its training corpus is dominated by Node.js code from before these APIs were universal—require('url'), querystring.parse(), express middleware patterns, axios with custom timeout wrappers, multer for form parsing. Those patterns are statistically dominant in what the model learned from. They’re what it reaches.

The gap between what the model defaults to and what the platform already provides is where most of the output token cost lives.

The Magnitude, by Pattern

I’ve been estimating the token economics of this as I go. These are approximate—based on the actual length of the patterns, not from a formal study—but the ratios are consistent enough to be useful.

Query parameter parsing

// model default—manual parsing (~140 tokens)
const parts = rawUrl.split('?');
const pairs = parts[1] ? parts[1].split('&') : [];
const params = {};
pairs.forEach(p => {
	const [k, v] = p.split('=');
	params[decodeURIComponent(k)] = decodeURIComponent(v);
});

// Web API (~12 tokens)
const params = Object.fromEntries(new URL(rawUrl).searchParams);

Roughly 140 tokens versus 12. About 90% reduction, per occurrence. The manual version also silently fails on malformed keys, silently drops all but the last value for repeated parameters, and is a prototype pollution vector if the key is __proto__. The native version handles all of it by specification.

Form data

// model default—per-field state (~200+ tokens for a 3-field form)
const [name, setName] = useState('');
const [email, setEmail] = useState('');
const [role, setRole] = useState('');
const handleChange = (e) =>
	setFields({ ...fields, [e.target.name]: e.target.value });

// Web API (~14 tokens)
const data = Object.fromEntries(new FormData(event.target));

The model will generate state tracking and change handlers for every field. The native version ingests the entire form in one call. Roughly 200–250 tokens versus 14, depending on field count—and the native version scales to twenty fields at the same cost.

Fetch lifecycle and cancellation

// model default (~90 tokens)
let timer;
const controller = new AbortController();
timer = setTimeout(() => controller.abort(), 5000);
try {
	const res = await fetch(url, { signal: controller.signal });
} finally {
	clearTimeout(timer);
}

// Web API (~12 tokens)
const res = await fetch(url, { signal: AbortSignal.timeout(5000) });

The manual version leaks timers if the finally path is missed during refactoring. The native version has no lifecycle to manage.

Parallel async with failure isolation

// model default (~100 tokens)
let anyFailed = false;
const results = await Promise.all(
	tasks.map(t => t.catch(e => { anyFailed = true; return null; }))
);
if (anyFailed) { /* now what? */ }

// Web API (~10 tokens)
const results = await Promise.allSettled(tasks);

Promise.allSettled() returns a structured result per task with .status of "fulfilled" or "rejected" and the corresponding value or reason. The manual version loses the error detail and invents a new ad hoc status convention on every use.

UI components

// model default—custom modal (~250 tokens of JS lifecycle management)
const [isOpen, setIsOpen] = useState(false);
useEffect(() => {
	if (isOpen) document.body.style.overflow = 'hidden';
	return () => { document.body.style.overflow = ''; };
}, [isOpen]);
// ... aria attributes, keyboard trap, backdrop click handler ...

// semantic HTML (~25 tokens)
<dialog ref={ref}>...</dialog>
// browser handles focus trap, Escape key, accessibility tree, backdrop

<dialog> has been supported across all major browsers since 2022. <details>/<summary> for accordions, native <form> constraint validation (required, type="email", pattern, minlength)—these are not obscure. The model reaches for JavaScript implementations because that’s what’s in its training data. It will keep doing this until directed otherwise.

A complete Deno request handler

The compound effect is where this becomes substantial. A Deno handler that parses request params, reads a form body, queries a database, and returns a response—written in the model’s default style—runs to 400–600 output tokens for the boilerplate alone, before any application logic. The same handler written with native APIs runs to 60–90 tokens. That’s not a marginal improvement.

// native Web APIs throughout (~70 tokens of infrastructure)
export async function handler(request) {
	const { searchParams } = new URL(request.url);
	const tenantId = searchParams.get('tenant');
	const data = Object.fromEntries(new FormData(await request.formData()));
	const result = await db.query(`
SELECT id, name
FROM records
WHERE tenant_id = ?
AND active = 1
`).bind(tenantId).first();
	return Response.json(result);
}

Security and Reliability as Structural Outcomes

This is worth naming directly rather than leaving as a footnote. Moving to native APIs doesn’t just reduce token cost—it eliminates categories of bugs.

Manual query string parsing with params[key] = value is a prototype pollution vector. Manual decodeURIComponent fails silently on % in certain positions. Custom setTimeout-based abort patterns leak when the cleanup path is skipped during refactoring. Custom form state tracking creates consistency bugs when a field is added but the handler isn’t updated. Homemade modal focus management routinely breaks keyboard navigation and screen readers.

The native implementations are spec-compliant. They’ve been tested against every edge case that exists in real web traffic. The Web Platform Tests suite runs tens of thousands of interoperability tests against each browser and runtime. URLSearchParams handles + encoding, repeated parameters, empty values, and UTF-8 edge cases correctly because it was written to the spec that defines what correct means. The model’s hand-rolled equivalent handles whatever the author thought of that day.

This is not a minor reliability improvement. It’s the difference between code that was implemented once by the person who wrote the spec versus code that was written from memory by a pattern-matching system trained on a corpus full of implementations that got it partly wrong.

What Comments Are Actually Doing

I’d thought of comments as documentation—useful for humans, neutral for LLMs. Research from MITRE published in June 2025 (Sabetto et al., tested across Claude, GPT-4, Llama, and Mixtral) changed that. Comments aren’t neutral. Models follow comment intent even when it contradicts the code. Inaccurate comments—comments that describe what the code used to do before a refactor—actively degraded LLM comprehension below the no-comment baseline. Worse than silence.

A stale comment isn’t harmless. It’s misinformation with authority. When a model keeps returning to a pattern I’ve moved away from, a stale comment near that code is a real candidate for why.

What comments are worth—what actually carries useful information—is design intent. Constraints. Why this function doesn’t catch its own errors. Why the SQL filters at the database level instead of in application code. What must not change when this is refactored. The reason for a non-obvious choice. That’s signal. “Loop over items” above items.forEach() is noise, and adds tokens with no return.

ACL 2024 work on comment augmentation supports the other direction: models trained on code with comments outperform models trained on uncommented code. Comments are a semantic bridge. At inference time they still carry signal, so the content of that signal matters.

The Formatting Question, Correctly Weighted

There is a real finding here. Pan, Sun et al. (“The Hidden Cost of Readability,” August 2025) measured input token overhead from formatting across tens of thousands of source files. Removing indentation, blank lines, and alignment whitespace reduced input token counts by an average of 24.5% with essentially no accuracy change for Claude or GPT-4.

That’s the input side, and it’s real. The tractable individual choices—no alignment whitespace, SQL ex-dented to the left margin, no blank lines inside function bodies—aggregate to roughly 5–10% input savings under typical JS conditions.

But input tokens cost one-third to one-fifth what output tokens cost. And the output savings from native APIs are not 5–10%—they’re 85–92% per pattern, compounding across every occurrence. The formatting work is worth doing. It is not the main event.

My preference for ex-dented SQL has a sound technical rationale: the model’s SQL training data is predominantly left-aligned, so matching that distribution makes sense. Whether it measurably improves accuracy I can’t point to a controlled JavaScript study for. It looks right to me, and the argument is sound enough.

What I’m Putting in Prompts [And Working Through]

The mechanism that actually changes model output is an explicit directive named at the start of the session. General style guidance produces marginal improvement—Wang et al. (ACM, 2024–2025) found this in a study of style-aware prompting. What works better is naming specific APIs explicitly, making the correct answer available before the model reaches for its training-data default.

Here’s what I’m actively working on. Note the regular use of DO THIS and NOT THAT–these work best together. (This works by constraining the probability space before generation, and is a recurring suggestion you can see across the examples described here.)

use Web APIs natively: URL, URLSearchParams, FormData, AbortController, fetch, Headers, Request, Response, Promise.allSettled(), Promise.any() use semantic HTML: <dialog>, <details>, <form> with native constraint validation. Do not implement in JavaScript what the browser or Deno runtime provides natively

Combined with comment discipline:

Comments state design constraints, invariants, and why. Not what the code does. Do not write comments that restate what the next line does.

The native API directive is the one that produces the most visible difference in output quality and cost.

Where This Lands

The core finding is structural, not a tip. Deno made the choice to implement the Web API surface natively, creating a single consistent set of abstractions that work identically in the browser and on the server. That surface solves—correctly, securely, and for free—a large category of problems that LLMs are currently solving again from scratch, badly, every generation, at 85–92% more token cost than necessary.

The comment findings matter because the model treats them as authoritative input, not metadata. Stale comments produce actively wrong output. Accurate design-intent comments constrain generation in useful directions.

The formatting findings are real and worth applying. They are secondary to the API question.

What’s striking to me is that the biggest lever here—the one that produces 7–10× output token reduction on infrastructure code and eliminates whole categories of security and reliability issues simultaneously—is not a new coding technique. It’s using what the platform already built. The friction is that the model doesn’t know to use it unless you say so. Once you do, it’s consistent about it. The model doesn't know what your runtime already ships. Someone has to—and that's the entire reason you hire professionals instead of just running the model.


Sources


This is what I’m finding in my own workflow. All of the token estimates above are early approximations from direct observation, not from published studies. The directional findings are highly consistent. Your specific numbers will vary with your codebase so test it to see what really works for you, your work and your team. jimmont.com