惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

让小产品的独立变现更简单 - ezindie.com
让小产品的独立变现更简单 - ezindie.com
WordPress大学
WordPress大学
人人都是产品经理
人人都是产品经理
Engineering at Meta
Engineering at Meta
小众软件
小众软件
I
InfoQ
有赞技术团队
有赞技术团队
Cyber Security Advisories - MS-ISAC
Cyber Security Advisories - MS-ISAC
Martin Fowler
Martin Fowler
月光博客
月光博客
雷峰网
雷峰网
aimingoo的专栏
aimingoo的专栏
云风的 BLOG
云风的 BLOG
Last Week in AI
Last Week in AI
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻
S
SegmentFault 最新的问题
The GitHub Blog
The GitHub Blog
Y
Y Combinator Blog
V
Visual Studio Blog
博客园 - 叶小钗
钛媒体:引领未来商业与生活新知
钛媒体:引领未来商业与生活新知
GbyAI
GbyAI
P
Proofpoint News Feed
Apple Machine Learning Research
Apple Machine Learning Research

OfficeChai

These Are The 10 Cheapest AI Models In The World [June 2026] 18 Best AI Tools For English Speaking (With Examples) [2026] AI Impact? Vacancy Rates For US Office Properties Are Now Highest Since The 2008 Crisis KPMG Pulls Report Praising AI After It Was Found To Have Fake AI-Generated Citations India's Sarvam Raises $234 Million At $1.5 Billion Valuation After SpaceX Stock Pops 20%, Musk Has Made More Money In The Last 24 Hours Than Warren Buffett Made In His Entire Career OfficeChai Nobody Is Using AI Better Than Meta: NVIDIA CEO Jensen Huang 21 Best AI Tools For Animation (With Examples) [2026] 22 Best AI Tools For Architecture (With Examples) [2026] Datacenter Construction Spending Has Eclipsed Public Transportation Spending In The US China Scraps 12,000 Degree Courses, Mainly In Arts And Humanities, To Prepare For AI Age OfficeChai There Is No Job Loss With AI: David Friedberg Loop Between Human Capital And "Token Capital" Will Be The New IP For Firms, Says Satya Nadella How to Reduce Dependency on Key Employees 8 Google Index Checker Use Cases Beyond New Blog Posts Memory Squeeze? Smartphone Purchases Are Down Globally 21 Best AI Tools For Accounting (With Examples) [2026] AI For Voice Generation: 22 Best Options (With Examples) [2026] These Are The Most Popular Image Generation Models On OpenRouter [June 2026] Search Traffic For Websites Is Down 25% Over The Last Year Because Of AI: a16z Data Agentic Coding Has Led To A 50% Increase In Number Of Apps, But Most Are Finding Very Few Users: SimilarWeb Data OpenRouter Launches Fusion API, Which Uses A Combination Of Models To Achieve Fable-Like Performance At Half The Price Dario Amodei Refused To De-Deploy Or Fix Vulnerabilities In Fable Before US Export Controls, Says David Sacks 23 Best AI Tools For Notes Making (With Examples) [2026] 16 Best AI Tools For Astrology (With Examples) [2026] How Jensen Huang Once Had To Ask SEGA's CEO To Pay NVIDIA For A Technology That Didn't Work ChatGPT Already Has 11% Of The Search Market: OpenAI CFO Sarah Friar SpaceX Has Now Launched More Satellites Than Rest Of Humanity Combined Across History
Law Professors Prefer AI Answers Over Those Of Their Peer...
OfficeChai Team · 2026-06-17 · via OfficeChai

AI is fast eclipsing the abilities of the top people in some of the highest-paid professions.

A new study by researchers from Stanford and other leading U.S. law schools has delivered striking evidence of this shift in one of the most demanding professional domains: legal education. Titled “Law Professors Prefer AI Over Peer Answers,” the paper finds that when law professors were asked to blindly choose between short-answer responses written by their colleagues and those generated by large language models (LLMs), they overwhelmingly preferred the AI versions.

The study, published May 27, 2026, involved sixteen contracts law professors from fourteen U.S. law schools who all teach from the same casebook. Participants first created 40 representative office-hours-style questions across categories like case recall, doctrine, hypotheticals, and policy. They then wrote their own answers and judged 2,918 anonymized pairwise comparisons between human and LLM responses.

Clear Preference for AI

Professors rated responses from Google’s Gemini 2.5 Pro at a 75.92% win rate against human instructors, while NotebookLM (a retrieval-augmented version grounded in the casebook) achieved 74.75%. The models performed on par with the strongest human participants, and in some analyses, even outperformed every instructor. Every single judge in the study preferred LLM answers over peer responses on average, with a median LLM-preference rate of 75.81%.

Notably, the advantage held across all question types—including complex hypotheticals and policy questions that require nuanced judgment rather than rote recall. AI responses were also flagged as pedagogically harmful far less often (3.53% pooled rate) compared to professor-written answers (12.06% average).

The researchers went further by engineering textual features such as length, clarity, structure, and pedagogical support to test whether surface-level polish explained the results. It didn’t. LLMs consistently outperformed predictions based on these features alone, suggesting the advantage stems from substantive reasoning quality.

Shared Professional Standards

To determine whether this reflected genuine alignment with expert standards or mere stylistic appeal, the team analyzed inter-judge agreement on overlapping trials. Agreement exceeded what would be expected from purely idiosyncratic preferences, indicating that LLMs were capturing latent professional norms that the professors themselves endorse.

Using an “LLM-as-judge” framework validated against human evaluators, the researchers extended the ranking to newer models. Claude Opus 4.7 topped the list, followed by other frontier systems. All outperformed human instructors. Reasoning-focused variants, such as Gemini 2.5 Flash with thinking budget, significantly outperformed non-reasoning counterparts.

Implications for Legal Education and Beyond

The findings challenge assumptions about AI’s limitations in high-judgment fields. While many prior evaluations focused on objective accuracy, this study tested AI against the subjective but shared standards of expert practitioners—the very essence of legal training.

For law schools facing instructor capacity constraints, the results point to a practical opportunity: always-available AI tutors that can deliver high-quality short answers aligned with professional expectations. The authors suggest implementations with clear guardrails, citations to source material, and escalation paths to human faculty.

The paper also highlights a curious detail: stock Gemini 2.5 Pro often outperformed RAG-grounded variants (including a commercial AI tutor built on the same base model), raising questions about context dilution in long-document retrieval.

As AI capabilities continue to advance rapidly, this research underscores a broader trend. In domains where success depends on reasoned judgment rather than single ground truths, frontier models are not just matching experts—they are frequently preferred by them. For businesses, technologists, and educators, the message is clear: the integration of AI into professional knowledge work is accelerating, even in fields long considered resistant to automation.