惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

Cyberwarzone
Cyberwarzone
Vercel News
Vercel News
freeCodeCamp Programming Tutorials: Python, JavaScript, Git & More
奇客Solidot–传递最新科技情报
奇客Solidot–传递最新科技情报
aimingoo的专栏
aimingoo的专栏
B
Blog RSS Feed
A
About on SuperTechFans
T
The Blog of Author Tim Ferriss
爱范儿
爱范儿
腾讯CDC
S
SegmentFault 最新的问题
Exploit-DB.com RSS Feed
Exploit-DB.com RSS Feed
The Hacker News
The Hacker News
J
Java Code Geeks
大猫的无限游戏
大猫的无限游戏
B
Blog
IT之家
IT之家
Spread Privacy
Spread Privacy
K
KPMG report finds enterprise disconnect between AI and its ROI | CIO
C
Cisco Blogs
Recent Announcements
Recent Announcements
H
Hacker News: Front Page
AI
AI
I
InfoQ
H
Heimdal Security Blog
T
Threatpost
Cisco Talos Blog
Cisco Talos Blog
Threat Intelligence Blog | Flashpoint
Threat Intelligence Blog | Flashpoint
I
Intezer
W
WeLiveSecurity
SecWiki News
SecWiki News
MongoDB | Blog
MongoDB | Blog
宝玉的分享
宝玉的分享
博客园 - 【当耐特】
云风的 BLOG
云风的 BLOG
T
Threat Research - Cisco Blogs
V2EX - 技术
V2EX - 技术
N
News and Events Feed by Topic
cs.CV updates on arXiv.org
cs.CV updates on arXiv.org
O
OpenAI News
阮一峰的网络日志
阮一峰的网络日志
T
Troy Hunt's Blog
www.infosecurity-magazine.com
www.infosecurity-magazine.com
博客园 - 司徒正美
Apple Machine Learning Research
Apple Machine Learning Research
雷峰网
雷峰网
T
Tor Project blog
有赞技术团队
有赞技术团队
Schneier on Security
Schneier on Security
Last Week in AI
Last Week in AI

Hacker News - Newest: "AI"

AI can't read an investor deck AI as an attorney? Student uses ChatGPT, Gemini to sue UW over alleged racial discrimination Hacking MCP Servers in AI Systems – The Rug Pull: Tool Changes After Approval GitHub - MeepCastana/KubeezCut: Free Web based video editor GitHub - GenAI-Gurus/awesome-eu-ai-act: Curated tools, official sources, OSS, templates, and guides for EU AI Act compliance. Can AI judge journalism? A Thiel-backed startup says yes, even if it risks chilling whistleblowers Coming soon: 10 Things That Matter in AI Right Now DARPA built an AI to fact-check enemy weapons claims What explains heterogeneity in AI adoption? When AI Meets Muscle: Context-Aware Electrical Stimulation Promises a New Way to Guide Human Movements - Department of Computer Science AI Changed How We Build. It Did Not Change What Matters. Linux rules on using AI-generated code - Copilot is OK, but humans must take 'full responsibility for the… Meta spins up AI version of Mark Zuckerberg to engage with employees Code Mode: Let Your AI Write Programs, Not Just Call Tools | TanStack Blog GitHub - Delavalom/graft: Go framework for building AI agents. Type-safe tools, multi-provider (OpenAI, Anthropic, Gemini, Bedrock), zero vendor SDKs. India's TCS tops estimates, says new AI models did not dent services demand Gen Z's fading AI hype Strong feeling: we are in a folded AI reality GitHub - machinarii/total-recall-catalog: A reference catalog of latest knowledge retrieval, memory & RAG systems GitHub - mensfeld/code-on-incus: Give each AI agent its own isolated machine with root, Docker, and systemd. Active defense detects and stops threats automatically.. Quantization, LoRA, and the 8% Problem: Benchmarking Local LLMs for Production AI Iran war: We spoke to the man making Lego-style AI videos that experts say are powerful propaganda Powell, Bessent discussed Anthropic's Mythos AI cyber threat with major U.S. banks GitHub - immartian/bellamem: Persistent belief-graph memory for AI agents. Retrieves decisive context by importance — not recency, not RAG, not /compact. recursive-mode: The Repo-Native Operating System for AI Engineering After the attack on Sam Altman's home, will AI CEO's go on the offensive? The biggest advance in AI since the LLM Opus 4.6 vs GPT 5.4 One Prompt Unity World Generation Test “AI polls” are fake polls Client Challenge Can AI be a 'child of God'? Inside Anthropic's meeting with Christian leaders How to Switch AI Chatbots and Why You Might Want To GitHub - MattMessinger1/agentic_refund_guardrail: Safe refund policy layer for AI agents — Python + TypeScript. Same behavior, shared tests. Adam/papers/emergent_values_whitepaper.md at master · strangeadvancedmarketing/Adam Ask HN: How do you stop playing 20 questions with your AI coding tools How far can automation and AI support psychotherapy? - @theU GitHub - stagas/rtdiff: realtime git diff gui and AI-assisted commits A Mac Studio for Local AI — 6 Months Later A History of the Early Years of AI at the University of Edinburgh Why AI Coding Tools Still Feel Stuck on Localhost MSN AI Datacenters Are Becoming Strategic Targets twitter.com Penn Researchers Use AI to Surface Unreported GLP-1 Side Effects in Reddit Posts Show HN: MoodSense AI (ML and FastAPI and Gradio, Deployed on Hugging Face) Moodsense Ai - a Hugging Face Space by aman179102 AI models are terrible at betting on soccer—especially xAI Grok GitHub - xialeistudio/echoic GitHub - HimashaHerath/github-dev-wrapped: AI-powered weekly GitHub activity reports deployed to GitHub Pages GitHub - alejandrobalderas/claude-code-from-source: Architecture, patterns & internals of Anthropic's AI coding agent — reverse-engineered from source maps AI and Tech brief: Ireland ascendant GitHub - Titovilal/context0: Context0 - Never Surrender Training for a Marathon with an AI Coach: What Worked and What Didn't Cyber Pulse: Agentic Intel - Apps on Google Play I Built an AI PR Reviewer That Catches Bugs by Not Looking for Bugs Gen Z workers are so fearful AI will take their job they’re intentionally sabotaging their company’s AI rollout | Fortune How AI Is Reimagining the Game of Golf–For Both Players and Courses GitHub - nattergabriel/reseed: A CLI tool for managing and distributing agent skills across projects Is SVG the final frontier? My AI workflow evolved from prompts to a near-autonomous workflow MLSharp Help - 3DGS Viewer & Generator I put my cognitive field based AI's runtime on GitHub Is Numble the first AI-proof game? A3: Kubernetes for autonomous AI agent fleets | Emergent Principles Deepali Vyas ("The Elite Recruiter") GitHub - msmarkgu/RelayFreeLLM: A restful API designed to route user prompts to various AI model providers. Unionized ProPublica staff are on strike over AI, layoffs, and wages Unleashing the Advantage of Quantum AI We're heading for an AI-fueled 'dementia crisis,' brain scientist warns The AI-Assisted Breach of Mexico's Government Infrastructure [pdf] GitHub - stef41/lmscan: 🔍 Detect AI-generated text and fingerprint which LLM wrote it. Open-source GPTZero alternative. Zero dependencies, works offline. MSN GitHub - visionscaper/collabmem: Enabling long-term collaboration with Agentic AI - building up episodic and world model memory over time with in-context awareness We gave an AI a 3 year retail lease in SF and asked it to make a profit | Andon Labs AI Code is Hollowing Out Open Source, and Maintainers are Looking the Other Way What leaked "SteamGPT" files could mean for the PC gaming platform's use of AI AI is the boss at this retail store. What could go wrong? GitHub - Wuzu11517/agentic-proxy: Local proxy meant to help reduce With Drones, Geophysics and ArtificiaI Intelligence, Researchers Prepare to Do Battle Against Land Mines A Single Operator, Two AI Platforms, Nine Government Agencies: The Full Technical Report 在 Steam 上购买 FriedrichAI: Offline AI 立省 10% GitHub - inevolin/resume-cli: Hit Claude usage limits? Resume any AI coding session elsewhere. Switch tools at zero friction. GitHub - atripati/ark: AI Runtime Kernel — a context operating system for AI agents. Eliminates tool bloat, loads only what’s needed, and gives LLMs their reasoning space back. How to Build a Secure AI PR Reviewer with Claude, GitHub Actions, and JavaScript This Startup Wants You to Pay Up to Talk With AI Versions of Human Experts Intel Arc Pro B70 Brings 32GB VRAM to Local AI for $949 WordPress 7.0: The Good, the AI, and the Still Missing AI on the couch: Anthropic gives Claude 20 hours of psychiatry IatroBench: Pre-Registered Evidence of Iatrogenic Harm from AI Safety Measures AI Agents Know About Supabase. They Don't Always Use It Right. The history and future of AI at Google, with Sundar Pichai Inside an AI‑enabled device code phishing campaign How Meta Used AI to Map Tribal Knowledge in Large-Scale Data Pipelines AI for Systems: Using LLMs to Optimize Database Query Execution Forecasting the Economic Effects of AI Introducing Tinker: Play with AI, bring your ideas to life AI sheds light on an ancient gaming mystery People really hate AI but not as much as Iran—or Democrats | Fortune What is an AI Product Engineer? Phoebe Gates wants her $185 million AI startup to succeed with 'no ties to my privilege or my last name': 'I have a chip on my shoulder' | Fortune
Tokenmaxxing and the search for AI metrics that matter
tonkkatonka · 2026-04-27 · via Hacker News - Newest: "AI"

You have 1 article left to read this month before you need to register a free LeadDev.com account.

Estimated reading time: 6 minutes

Key takeaways:

  • Token usage is the lines-of-code metric of the AI era – easy to measure, easy to game, and disconnected from actual productivity.
  • The best frameworks track cognitive delegation, not consumption.
  • Self-reporting works, but only where trust exists.

Some engineering organizations are measuring AI output via tokens burned, some compare engineers to executive chefs, and some rely on self-reporting.

When Meta’s internal tokenmaxxing leaderboard – ranking engineers by how many AI tokens they consumed – became public knowledge, it didn’t take long for the engineering community to react. The leaderboard has since been taken down, but high token usage is not a badge of honor just at Meta.

Across large organizations, there is enormous pressure to prove that the millions being spent on AI tooling are paying off. 

Your inbox, upgraded.

Receive weekly engineering insights to level up your leadership approach.

The easiest metric to game

Tokens burned has become the easiest number to point to. It’s objective, it’s automated, it scales across thousands of engineers, and it gives leadership a dashboard. It’s easy to measure. It’s also easy to game.

Measuring AI adoption and productivity gains is hard, and especially if you have to do it at an individual level. Token usage definitely isn’t the right metric. You could just use tokens to run your OpenClaw!” says Ankit Jain, founder of the Hangar, a community of senior engineers and engineering leaders focused on developer experience and solving productivity challenges at scale.

Tokenmaxxing is going back to the pre-DevOps Research and Assessment (DORA) and to ‘measuring lines of code’ era, he adds. DORA metrics are a set of five software delivery performance metrics that provide an effective way of measuring the outcomes of the software delivery process.

However, DORA metrics were never designed to evaluate individual developer output. They work as multiple metrics, not just one, Jain says. If you game one metric, it’s going to break down other metrics.

When organizational trust is low, managers reach for data to justify headcount and budget decisions, and individual productivity measurement follows. That’s how DORA gets misapplied as a personal scorecard, and the same dynamic is now producing tokenmaxxing leaderboards.

“Organizations need to figure out how much outcome AI tools are driving, and that’s a hard problem to measure,” Jain says. The solution, he argues, is combined metrics: “We have to come up with a set of metrics versus just one. Tokens used, though, has to be one of them.”

From line cook to executive chef: a skill-level-based framework

One of the more interesting attempts to move beyond token counting comes from a staff engineer at a FAANG company we spoke to, who has been building an AI adoption measurement framework at their organization.

The framework is inspired by Steve Yegge’s ‘executive chef’ model that compares software engineers to chefs in a restaurant kitchen. AI agents are engineers’ sous chefs, line cooks, and prep cooks. Just like executive chefs don’t do all the chopping themselves but decide what goes on the menu, develop the recipes, and taste everything before it’s served, engineers orchestrate agents and own the outcome.

They have adapted Yegge’s original nine levels to four, ranging from basic interactive tool use to orchestrating multiple autonomous agents. The levels do not measure which tools engineers use or how many tokens they burn, but rather the degree of cognitive work being delegated to AI over time, and how effectively.

“A developer at level one is using AI for quick one-shot queries. At level four, they’re writing detailed specifications, orchestrating multiple agents, and shifting quality assurance left toward spec review rather than code review,” they say.

The framework also revealed a counterintuitive proxy signal: as engineers progress through the levels, they file fewer bugs because they resolve issues inline rather than queue them. A declining bug backlog growth rate becomes a lagging indicator of genuine AI maturity – not something you’d normally think to instrument for, but meaningful once you know to look.

The data also surfaced a bigger problem: the level progression is not linear or obvious. Engineers at level one can’t easily imagine level four. There is a real adoption challenge that token dashboards detect but can’t diagnose: engineer resistance.

“There is a genuine pushback of ‘I don’t want my job to be an orchestrator of agents,’ which is an identity concern, not a capability one, and a much harder thing to measure or address. Any honest framework for AI adoption has to account for where an organization sits on that curve,” they say.

The opposite approach: self-reporting and trust

Emily Nakashima, SVP of engineering at Honeycomb, says there is a lot of fear among people at software companies (not just among engineers) when the question of measuring AI impact and productivity is raised. People worry about what that means for their jobs.

“We did a top-down founder memo last summer, saying that we really believe in this new technology, we really want people to spend time experimenting and learning about it, and that we should all try to 2x our impact with AI over the next year. The first question we got from engineers was ‘how are we going to measure this?’” she says.

Rather than building formal frameworks, Nakashima says they deliberately de-emphasized measurement, relying instead on self-reporting.

“I really worry about companies trying to 2x or 3x their token spend, because there are ways engineers can do that that return no value to the company. We actually get a lot of value out of self-reporting. For that to work well, you have to have a relatively high-trust organization. What engineers on my team give back in terms of self-reporting actually aligns pretty well with what’s seen in their work. When it’s working, you can see it, and these measurement questions go away a little bit.”

Going forward, Nakashima says she’d like to have an individual-level self-report and a manager-level self-report for the team on how much they have increased their impact with AI. She guesses that would be one of the most accurate or most valuable measurements. Honeycomb’s self-reporting approach requires a level of organizational trust that many large companies don’t have. 

Hard to measure, but it has to be measured

When trust is low, managers lean on hard numbers, and that’s how we end up with leaderboards. The alternative of not measuring at all isn’t viable either. Organizations are spending millions on AI tooling and need to know whether it’s working. Engineering leaders need signals to know where to invest in training, which teams need support, and where the productivity gains are happening. 

Before organizations can measure whether AI is making engineers more effective and where the skill gaps are, managers need to know whether they’re using it at all. Token usage is a good proxy for AI adoption, but it’s a terrible metric for AI productivity.

Angie Jones, former VP of engineering, AI tools, and enablement at Block, confirms that measuring token usage was helpful when they were early in their journey and she needed to track adoption.

“After adoption was clear, I threw it out as it did nothing to measure developer productivity. I wouldn’t be surprised if we see the opposite trend next year, aiming for efficient usage of tokens as opposed to celebrating burning them at expensive rates.”

LDX3 London 2026 agenda is live - See who is in the lineup

LondonJune 2 & 3, 2026

New activities added to LDX3 London 🎉

The tokenmaxxing backlash

The backlash to Meta’s tokenmaxxing leaderboard shows that the industry is aware that celebrating high token usage is not the right metric, especially as AI tooling costs continue to grow.

However, there is no consensus yet on what to measure instead. Jain suggests combining token usage with lines of code shipped and says some AI tools already provide data points around how much code was accepted. “Although ‘all the code was accepted’ is a very fuzzy measurement,” Jain admits.

That fuzziness is the state of things. The tooling to connect token spend to shipping outcomes doesn’t fully exist yet. The signals are noisy, and every framework requires tradeoffs between rigor and trust.

If Jones is right that the next phase will be about efficient token usage rather than maximum token usage, the leaderboards may not disappear – they’ll just measure something worth competing over.