惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

Stack Overflow Blog
Stack Overflow Blog
云风的 BLOG
云风的 BLOG
Cyber Security Advisories - MS-ISAC
Cyber Security Advisories - MS-ISAC
Recent Announcements
Recent Announcements
Microsoft Security Blog
Microsoft Security Blog
Microsoft Azure Blog
Microsoft Azure Blog
J
Java Code Geeks
D
DataBreaches.Net
U
Unit 42
P
Proofpoint News Feed
I
InfoQ
Apple Machine Learning Research
Apple Machine Learning Research
Google DeepMind News
Google DeepMind News
博客园 - Franky
博客园_首页
IT之家
IT之家
博客园 - 叶小钗
freeCodeCamp Programming Tutorials: Python, JavaScript, Git & More
酷 壳 – CoolShell
酷 壳 – CoolShell
博客园 - 【当耐特】
Hugging Face - Blog
Hugging Face - Blog
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻
让小产品的独立变现更简单 - ezindie.com
让小产品的独立变现更简单 - ezindie.com
阮一峰的网络日志
阮一峰的网络日志

Snorkel AI

Building AI-Native Systems for Federal Infrastructure: A Conversation with Rezaur Rahman Code World Models and AutoHarness for LLM Agents Benchtalks #1: Alex Shaw (Terminal-Bench, Harbor) – Building the Benchmark Factory Building FinQA: An Open RL Environment for Financial Reasoning Agents How Tool Discipline Let a 4B Model Outsmart a 235B Giant on Financial Tasks Coding agents don’t need to be perfect, they need to recover Closing the Evaluation Gap in Agentic AI SlopCodeBench: Measuring Code Erosion as Agents Iterate Introducing the Snorkel Agentic Coding Benchmark 2026: The year of environments Part V: Future Direction and Emerging Trends in Rubric-Based AI Evaluation The self-critique paradox: Why AI verification fails where it’s needed most Chat With the Terminal-Bench Team | Snorkel AI Terminal-Bench 2.0: Raising the bar for AI agent evaluation Snorkeling in RL environments Introducing SnorkelSpatial: A Benchmark for LLM Spatial Reasoning Scaling Trust: Rubrics in Snorkel's Quality Process Evaluating Multi-Agent Systems in Enterprise Tool Use Evaluating Coding Agents with Terminal-Bench 2.0 Parsing isn’t neutral: why evaluation choices matter The science of rubric design The right tool for the job: An A-Z of rubrics Data quality and rubrics: how to build trust in your models Building the benchmark: inside our agentic insurance underwriting dataset Evaluating AI agents for insurance underwriting LLM observability: key practices, tools, and challenges Anthropic Claude + AWS: revolutionizing pharma data analytics with Snorkel AI Data-centric development of an enterprise AI agent with Snorkel Building the data development platform for specialized AI LLM-as-a-judge for enterprises: evaluate model alignment at scale
Intelligence per watt: A new metric for AI’s future
11450pwpadmin · 2025-11-13 · via Snorkel AI

The AI community has been obsessed with bigger models and more data centers. But researchers at Stanford’s Hazy Research Lab are proposing we optimize for something entirely different.

They’ve introduced Intelligence per watt (IPW)—a new metric that fundamentally reframes how we should think about AI utilization in an era of exploding demand. Their paper breaks down the challenge and the opportunity, pointing us toward a compelling path forward for future research and innovation.

Efficiency is critical to meet ever-growing demand

Demand for AI computation is growing exponentially, with Google reporting an 8.1x increase in tokens processed per month from February 2024 to October 2025. However, the Hazy Research team also observes that internal ChatGPT telemetry data shows 77% of requests are practical tasks like writing emails or summarizing documents. In other words, for well over three fourths of real-world AI usage, we’re shipping routine queries–requests that could be answered accurately on the local device–to frontier-level models in datacenters. 

History offers a better path. From 1946-2009, computing efficiency doubled every 1.5 years, shifting workloads from mainframes to PCs. PCs won not through raw performance, but because efficiency improvements made computing capable enough within personal device power constraints.

We’re at that same inflection point with AI inference. Can we get more of our needs met on the edge, where power efficiency is greater and the absolute maximum AI reasoning capabilities are unnecessary? Can the exponential growth in demand for AI be met more effectively through better leverage of the devices in our pockets and backpacks? The Hazy Research team says yes!

Hazy Research’s intelligence per watt

The Hazy Research team defined IPW elegantly:

IPW = (mean accuracy across tasks) / (mean power draw during inference)

Their empirical study—20+ local models, diverse hardware, 1 million real-world queries—reveals three key findings:

  1. Local LMs accurately respond to 88.7% of single-turn queries, with accuracy improving 3.1× from 2023-2025
  2. Local accelerators have significant efficiency headroom—the M4 Max achieves 1.5× lower IPW than NVIDIA B200 for the same model
  3. Intelligence efficiency has improved 5.3× over the past two years through combined model and hardware advances

Snorkel AI’s contribution to the IPW initiative

At Snorkel AI, we’ve built benchmarks to evaluate frontier LLMs across expert-level, domain-specific tasks using our Expert Data-as-a-Service—powered by a global network of specialists across thousands of domains.

We’re excited to contribute these specialized datasets to Hazy Research Lab’s Intelligence Per Watt initiative. While their foundational work focused on general chat and reasoning, real-world deployment demands domain-specific evaluation.

By combining Hazy Research’s IPW measurement framework with Snorkel’s industry-relevant benchmarks—spanning insurance underwriting, financial analysis, legal review, and PhD-level technical domains—we can drive an industry-wide shift in how we approach AI’s compute needs.

This partnership will answer critical questions: How efficiently can local models handle medical reasoning? What’s the IPW for regulatory compliance tasks? Can edge devices deliver expert-level performance within power budgets?

The path forward

Hazy Research’s Intelligence Per Watt metric should guide AI’s transition to the edge, just as performance-per-watt guided the mainframe-to-PC shift. They’re releasing a hardware-agnostic profiling harness to make IPW measurement systematic and accessible. 

The future of AI isn’t just bigger models—it’s smarter systems delivering the right intelligence, in the right place, with the right efficiency. Snorkel AI is proud to support this vision with specialized datasets that ensure IPW becomes an important consideration for real-world enterprise deployment.


Read the full paper here and check out hazyresearch.stanford.edu for more information about Stanford University’s Hazy Research Lab, headed by Snorkel AI cofounder Chris Ré. Learn more about Snorkel AI’s data-centric approach at snorkel.ai.