惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

WordPress大学
WordPress大学
J
Java Code Geeks
Martin Fowler
Martin Fowler
Microsoft Azure Blog
Microsoft Azure Blog
月光博客
月光博客
freeCodeCamp Programming Tutorials: Python, JavaScript, Git & More
人人都是产品经理
人人都是产品经理
有赞技术团队
有赞技术团队
爱范儿
爱范儿
Engineering at Meta
Engineering at Meta
GbyAI
GbyAI
博客园 - 【当耐特】
Y
Y Combinator Blog
Last Week in AI
Last Week in AI
MongoDB | Blog
MongoDB | Blog
G
Google Developers Blog
博客园 - 三生石上(FineUI控件)
Cyber Security Advisories - MS-ISAC
Cyber Security Advisories - MS-ISAC
大猫的无限游戏
大猫的无限游戏
罗磊的独立博客
The Cloudflare Blog
钛媒体:引领未来商业与生活新知
钛媒体:引领未来商业与生活新知
V
V2EX
博客园 - 司徒正美

Snorkel AI

Building AI-Native Systems for Federal Infrastructure: A Conversation with Rezaur Rahman Code World Models and AutoHarness for LLM Agents Benchtalks #1: Alex Shaw (Terminal-Bench, Harbor) – Building the Benchmark Factory Building FinQA: An Open RL Environment for Financial Reasoning Agents How Tool Discipline Let a 4B Model Outsmart a 235B Giant on Financial Tasks Coding agents don’t need to be perfect, they need to recover Closing the Evaluation Gap in Agentic AI SlopCodeBench: Measuring Code Erosion as Agents Iterate Introducing the Snorkel Agentic Coding Benchmark 2026: The year of environments Part V: Future Direction and Emerging Trends in Rubric-Based AI Evaluation The self-critique paradox: Why AI verification fails where it’s needed most Chat With the Terminal-Bench Team | Snorkel AI Terminal-Bench 2.0: Raising the bar for AI agent evaluation Snorkeling in RL environments Introducing SnorkelSpatial: A Benchmark for LLM Spatial Reasoning Scaling Trust: Rubrics in Snorkel's Quality Process Evaluating Multi-Agent Systems in Enterprise Tool Use Evaluating Coding Agents with Terminal-Bench 2.0 Parsing isn’t neutral: why evaluation choices matter The science of rubric design The right tool for the job: An A-Z of rubrics Data quality and rubrics: how to build trust in your models Building the benchmark: inside our agentic insurance underwriting dataset Evaluating AI agents for insurance underwriting LLM observability: key practices, tools, and challenges Anthropic Claude + AWS: revolutionizing pharma data analytics with Snorkel AI Data-centric development of an enterprise AI agent with Snorkel Building the data development platform for specialized AI LLM-as-a-judge for enterprises: evaluate model alignment at scale
Intelligence per watt: A new metric for AI’s future
11450pwpadmin · 2025-11-13 · via Snorkel AI

The AI community has been obsessed with bigger models and more data centers. But researchers at Stanford’s Hazy Research Lab are proposing we optimize for something entirely different.

They’ve introduced Intelligence per watt (IPW)—a new metric that fundamentally reframes how we should think about AI utilization in an era of exploding demand. Their paper breaks down the challenge and the opportunity, pointing us toward a compelling path forward for future research and innovation.

Efficiency is critical to meet ever-growing demand

Demand for AI computation is growing exponentially, with Google reporting an 8.1x increase in tokens processed per month from February 2024 to October 2025. However, the Hazy Research team also observes that internal ChatGPT telemetry data shows 77% of requests are practical tasks like writing emails or summarizing documents. In other words, for well over three fourths of real-world AI usage, we’re shipping routine queries–requests that could be answered accurately on the local device–to frontier-level models in datacenters. 

History offers a better path. From 1946-2009, computing efficiency doubled every 1.5 years, shifting workloads from mainframes to PCs. PCs won not through raw performance, but because efficiency improvements made computing capable enough within personal device power constraints.

We’re at that same inflection point with AI inference. Can we get more of our needs met on the edge, where power efficiency is greater and the absolute maximum AI reasoning capabilities are unnecessary? Can the exponential growth in demand for AI be met more effectively through better leverage of the devices in our pockets and backpacks? The Hazy Research team says yes!

Hazy Research’s intelligence per watt

The Hazy Research team defined IPW elegantly:

IPW = (mean accuracy across tasks) / (mean power draw during inference)

Their empirical study—20+ local models, diverse hardware, 1 million real-world queries—reveals three key findings:

  1. Local LMs accurately respond to 88.7% of single-turn queries, with accuracy improving 3.1× from 2023-2025
  2. Local accelerators have significant efficiency headroom—the M4 Max achieves 1.5× lower IPW than NVIDIA B200 for the same model
  3. Intelligence efficiency has improved 5.3× over the past two years through combined model and hardware advances

Snorkel AI’s contribution to the IPW initiative

At Snorkel AI, we’ve built benchmarks to evaluate frontier LLMs across expert-level, domain-specific tasks using our Expert Data-as-a-Service—powered by a global network of specialists across thousands of domains.

We’re excited to contribute these specialized datasets to Hazy Research Lab’s Intelligence Per Watt initiative. While their foundational work focused on general chat and reasoning, real-world deployment demands domain-specific evaluation.

By combining Hazy Research’s IPW measurement framework with Snorkel’s industry-relevant benchmarks—spanning insurance underwriting, financial analysis, legal review, and PhD-level technical domains—we can drive an industry-wide shift in how we approach AI’s compute needs.

This partnership will answer critical questions: How efficiently can local models handle medical reasoning? What’s the IPW for regulatory compliance tasks? Can edge devices deliver expert-level performance within power budgets?

The path forward

Hazy Research’s Intelligence Per Watt metric should guide AI’s transition to the edge, just as performance-per-watt guided the mainframe-to-PC shift. They’re releasing a hardware-agnostic profiling harness to make IPW measurement systematic and accessible. 

The future of AI isn’t just bigger models—it’s smarter systems delivering the right intelligence, in the right place, with the right efficiency. Snorkel AI is proud to support this vision with specialized datasets that ensure IPW becomes an important consideration for real-world enterprise deployment.


Read the full paper here and check out hazyresearch.stanford.edu for more information about Stanford University’s Hazy Research Lab, headed by Snorkel AI cofounder Chris Ré. Learn more about Snorkel AI’s data-centric approach at snorkel.ai.