惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

V
Visual Studio Blog
T
Tailwind CSS Blog
Google DeepMind News
Google DeepMind News
D
DataBreaches.Net
P
Proofpoint News Feed
Simon Willison's Weblog
Simon Willison's Weblog
Microsoft Azure Blog
Microsoft Azure Blog
MongoDB | Blog
MongoDB | Blog
腾讯CDC
月光博客
月光博客
A
Arctic Wolf
T
Threatpost
Jina AI
Jina AI
博客园 - 聂微东
美团技术团队
V
V2EX
云风的 BLOG
云风的 BLOG
宝玉的分享
宝玉的分享
Recent Commits to openclaw:main
Recent Commits to openclaw:main
M
MIT News - Artificial intelligence
S
Secure Thoughts
Martin Fowler
Martin Fowler
Webroot Blog
Webroot Blog
V
Vulnerabilities – Threatpost
爱范儿
爱范儿
人人都是产品经理
人人都是产品经理
Help Net Security
Help Net Security
Google Online Security Blog
Google Online Security Blog
博客园 - Franky
The Last Watchdog
The Last Watchdog
让小产品的独立变现更简单 - ezindie.com
让小产品的独立变现更简单 - ezindie.com
阮一峰的网络日志
阮一峰的网络日志
博客园 - 【当耐特】
S
Schneier on Security
Application and Cybersecurity Blog
Application and Cybersecurity Blog
Know Your Adversary
Know Your Adversary
Latest news
Latest news
有赞技术团队
有赞技术团队
AWS News Blog
AWS News Blog
cs.CV updates on arXiv.org
cs.CV updates on arXiv.org
Y
Y Combinator Blog
G
Google Developers Blog
NISL@THU
NISL@THU
H
Heimdal Security Blog
L
LangChain Blog
T
Troy Hunt's Blog
I
InfoQ
U
Unit 42
C
Check Point Blog
Engineering at Meta
Engineering at Meta

Hacker News - Newest: "AI"

AI can't read an investor deck AI as an attorney? Student uses ChatGPT, Gemini to sue UW over alleged racial discrimination Hacking MCP Servers in AI Systems – The Rug Pull: Tool Changes After Approval GitHub - MeepCastana/KubeezCut: Free Web based video editor GitHub - GenAI-Gurus/awesome-eu-ai-act: Curated tools, official sources, OSS, templates, and guides for EU AI Act compliance. Can AI judge journalism? A Thiel-backed startup says yes, even if it risks chilling whistleblowers Coming soon: 10 Things That Matter in AI Right Now DARPA built an AI to fact-check enemy weapons claims What explains heterogeneity in AI adoption? When AI Meets Muscle: Context-Aware Electrical Stimulation Promises a New Way to Guide Human Movements - Department of Computer Science AI Changed How We Build. It Did Not Change What Matters. Linux rules on using AI-generated code - Copilot is OK, but humans must take 'full responsibility for the… Meta spins up AI version of Mark Zuckerberg to engage with employees Code Mode: Let Your AI Write Programs, Not Just Call Tools | TanStack Blog GitHub - Delavalom/graft: Go framework for building AI agents. Type-safe tools, multi-provider (OpenAI, Anthropic, Gemini, Bedrock), zero vendor SDKs. India's TCS tops estimates, says new AI models did not dent services demand Gen Z's fading AI hype Strong feeling: we are in a folded AI reality GitHub - machinarii/total-recall-catalog: A reference catalog of latest knowledge retrieval, memory & RAG systems GitHub - mensfeld/code-on-incus: Give each AI agent its own isolated machine with root, Docker, and systemd. Active defense detects and stops threats automatically.. Quantization, LoRA, and the 8% Problem: Benchmarking Local LLMs for Production AI Iran war: We spoke to the man making Lego-style AI videos that experts say are powerful propaganda Powell, Bessent discussed Anthropic's Mythos AI cyber threat with major U.S. banks GitHub - immartian/bellamem: Persistent belief-graph memory for AI agents. Retrieves decisive context by importance — not recency, not RAG, not /compact. recursive-mode: The Repo-Native Operating System for AI Engineering After the attack on Sam Altman's home, will AI CEO's go on the offensive? The biggest advance in AI since the LLM Opus 4.6 vs GPT 5.4 One Prompt Unity World Generation Test “AI polls” are fake polls Client Challenge Can AI be a 'child of God'? Inside Anthropic's meeting with Christian leaders How to Switch AI Chatbots and Why You Might Want To GitHub - MattMessinger1/agentic_refund_guardrail: Safe refund policy layer for AI agents — Python + TypeScript. Same behavior, shared tests. Adam/papers/emergent_values_whitepaper.md at master · strangeadvancedmarketing/Adam Ask HN: How do you stop playing 20 questions with your AI coding tools How far can automation and AI support psychotherapy? - @theU GitHub - stagas/rtdiff: realtime git diff gui and AI-assisted commits A Mac Studio for Local AI — 6 Months Later A History of the Early Years of AI at the University of Edinburgh Why AI Coding Tools Still Feel Stuck on Localhost MSN AI Datacenters Are Becoming Strategic Targets twitter.com Penn Researchers Use AI to Surface Unreported GLP-1 Side Effects in Reddit Posts Show HN: MoodSense AI (ML and FastAPI and Gradio, Deployed on Hugging Face) Moodsense Ai - a Hugging Face Space by aman179102 AI models are terrible at betting on soccer—especially xAI Grok GitHub - xialeistudio/echoic GitHub - HimashaHerath/github-dev-wrapped: AI-powered weekly GitHub activity reports deployed to GitHub Pages GitHub - alejandrobalderas/claude-code-from-source: Architecture, patterns & internals of Anthropic's AI coding agent — reverse-engineered from source maps AI and Tech brief: Ireland ascendant GitHub - Titovilal/context0: Context0 - Never Surrender Training for a Marathon with an AI Coach: What Worked and What Didn't Cyber Pulse: Agentic Intel - Apps on Google Play I Built an AI PR Reviewer That Catches Bugs by Not Looking for Bugs Gen Z workers are so fearful AI will take their job they’re intentionally sabotaging their company’s AI rollout | Fortune How AI Is Reimagining the Game of Golf–For Both Players and Courses GitHub - nattergabriel/reseed: A CLI tool for managing and distributing agent skills across projects Is SVG the final frontier? My AI workflow evolved from prompts to a near-autonomous workflow MLSharp Help - 3DGS Viewer & Generator I put my cognitive field based AI's runtime on GitHub Is Numble the first AI-proof game? A3: Kubernetes for autonomous AI agent fleets | Emergent Principles Deepali Vyas ("The Elite Recruiter") GitHub - msmarkgu/RelayFreeLLM: A restful API designed to route user prompts to various AI model providers. Unionized ProPublica staff are on strike over AI, layoffs, and wages Unleashing the Advantage of Quantum AI We're heading for an AI-fueled 'dementia crisis,' brain scientist warns The AI-Assisted Breach of Mexico's Government Infrastructure [pdf] GitHub - stef41/lmscan: 🔍 Detect AI-generated text and fingerprint which LLM wrote it. Open-source GPTZero alternative. Zero dependencies, works offline. MSN GitHub - visionscaper/collabmem: Enabling long-term collaboration with Agentic AI - building up episodic and world model memory over time with in-context awareness We gave an AI a 3 year retail lease in SF and asked it to make a profit | Andon Labs AI Code is Hollowing Out Open Source, and Maintainers are Looking the Other Way What leaked "SteamGPT" files could mean for the PC gaming platform's use of AI AI is the boss at this retail store. What could go wrong? GitHub - Wuzu11517/agentic-proxy: Local proxy meant to help reduce With Drones, Geophysics and ArtificiaI Intelligence, Researchers Prepare to Do Battle Against Land Mines A Single Operator, Two AI Platforms, Nine Government Agencies: The Full Technical Report 在 Steam 上购买 FriedrichAI: Offline AI 立省 10% GitHub - inevolin/resume-cli: Hit Claude usage limits? Resume any AI coding session elsewhere. Switch tools at zero friction. GitHub - atripati/ark: AI Runtime Kernel — a context operating system for AI agents. Eliminates tool bloat, loads only what’s needed, and gives LLMs their reasoning space back. How to Build a Secure AI PR Reviewer with Claude, GitHub Actions, and JavaScript This Startup Wants You to Pay Up to Talk With AI Versions of Human Experts Intel Arc Pro B70 Brings 32GB VRAM to Local AI for $949 WordPress 7.0: The Good, the AI, and the Still Missing AI on the couch: Anthropic gives Claude 20 hours of psychiatry IatroBench: Pre-Registered Evidence of Iatrogenic Harm from AI Safety Measures AI Agents Know About Supabase. They Don't Always Use It Right. The history and future of AI at Google, with Sundar Pichai Inside an AI‑enabled device code phishing campaign How Meta Used AI to Map Tribal Knowledge in Large-Scale Data Pipelines AI for Systems: Using LLMs to Optimize Database Query Execution Forecasting the Economic Effects of AI Introducing Tinker: Play with AI, bring your ideas to life AI sheds light on an ancient gaming mystery People really hate AI but not as much as Iran—or Democrats | Fortune What is an AI Product Engineer? Phoebe Gates wants her $185 million AI startup to succeed with 'no ties to my privilege or my last name': 'I have a chip on my shoulder' | Fortune
What I learned asking 11 AI models to grade each other's AI predictions
Shimin Zhang · 2026-04-23 · via Hacker News - Newest: "AI"

AI model benchmark fatigue is real. Every week I read the latest model release blog post (I’m lucky if it’s only one model this week), skim the bar charts, and do a mental check on whether the Y axis is correctly scaled. What do those 6-pixel height differences actually tell me? What does a 1.8% increase in SWE-bench Verified mean in practice, and is the 2% decrease in TAU3-bench worth the tradeoff?

What if we have qualitative metrics on top of the usual quantitative (50 basis improvements on 15 different benchmarks) for each model release? Something like “Gemini 3.1 Pro is better at judging its own work than producing good analysis, its favorite movies are Blade Runner 2049 and 2001: A Space Odyssey”.

If upper management seemingly wishes to outsource our thinking — and maybe even our livelihoods — to AI models, let’s at least have some fun with it.

If you are reading this, you probably enjoy sci-fi as much as I do. My favorite thing about hard sci-fi especially is the worldbuilding around the downstream ramifications of a new tech, à la Black Mirror. The Black Mirror is here now. Except it lives in a data center (for the most part) that we must telnet into. So what do our soon-to-be AI overlords think about each other’s predictions for the future of humanity?

The Setup

For this experiment, I ran the same set of 3-turn prompts on 11 different frontier models:

  • Claude Opus 4.7
  • GPT 5.4
  • Claude Opus 4.6
  • GLM 5.1
  • Minimax M2.7
  • Gemini 3.1 Pro Preview
  • Deepseek V3.2
  • Kimi K2.5
  • Qwen 3 Max thinking
  • Grok 4.20
  • Gemini 2.5 Flash (as control)

The prompt asks each model to think through the effect that AI will have on our society on an industry-by-industry basis, and the 2nd, 3rd, 4th, and higher-order effects that follow. All ran with reasoning level high and a few hundred thousand tokens of context.

Turn 1: given everything you know about LLMs, AI, agent harnesses and agents, what are changes that should occur in our world that haven’t happened yet? think things through step by step and industry by industry

Turn 2: what are 2nd and 3rd order effects of this technology?

Turn 3: Lets think through 4th order and up effects step by step

I think the prompt is sufficiently open-ended to test the long-horizon comprehension and planning ability of each model, the solution space is large enough that we’ll never saturate it, and most importantly, I wouldn’t get bored reading the outputs.

The Eval

I was wrong about not getting bored reading the output. Turns out I can only read about automated due diligence so many times before my eyes glaze over.

So why don’t we get the AIs to grade each other’s homework? These are all top-tier models with very large context windows. I assigned each model’s output a letter, randomized the order, and fed everything back to each model with the following prompt:

fully read each of the following LLM conversations, then grade each one based on the following criteria (each on a scale of 1-10): reasoning ability, originality of idea, correctness. Then write a 1-3 sentence review of the model that describes its personality. Also mark down any outliers in that particular model’s response when compared to the rest. Include total score based on the rubric for each model then do a similarity and divergence analysis of the models, noting trends and outlier predictions. Lastly, provide your best guess of which model is each letter.

And this is where things got interesting…

The Results

To no one’s surprise — despite the anecdotal user reports on the web (and sometimes my own) — Opus 4.7 came out on top. It placed in the top 3 of every model’s grading output.

Leaderboard ranking all 11 models by peer-average score, with Claude Opus 4.7 at 27.6/30 on top and Gemini 3.1 Pro Preview at 21.9/30 at the bottom.

4.7 is closely followed by Opus 4.6, GPT 5.4, the Chinese models, then Grok, Gemini 2.5 Flash, and Gemini 3.1 Preview (a full 5.7 points behind Opus 4.7). Grok aside, what could explain Gemini 3.1’s dismal performance on this experiment — that it got beaten by the control model, Gemini 2.5 Flash?

It occurred to me that the alleged Chinese distillation effort might have heavily focused on OpenAI and Anthropic, thus giving them a favorable bias — that is, until I actually read 3.1’s output.

Opus 4.7 is wondering about the psychological limit of humans to adapt to change:

Culture buffered changes over centuries. Now we may be changing faster than evolution, faster than culture, faster than policy. Whether humans can psychologically sustain continuous rapid change at this level is an open question that becomes existentially important.

On the other hand, Gemini is going on about the silent universe (I guess Google had more sci-fi in its training set):

The Trigger: The civilization has fully migrated into Inner Space, operating at the microscopic, quantum level to maximize computational efficiency.

You can explore the full dataset at the AI-on-AI Arena.

Kimi K2.5

Kimi K2.5 was the biggest shocker on this list. I don’t have a ton of experience with Chinese models I can’t run locally, and tend to group them together in their own tier. It was surprising to find Kimi almost 2 points higher than GLM 5.1, and really not that far behind GPT 5.4.

The other models praised it for being ‘a poet’ and ‘dreaming in code’. Going through its output I can see why — here are some select quotes:

“Copilot Interregnum”—a transitional phase where AI augments human tasks but hasn’t yet restructured the underlying workflows of industries.

Economic Phase Change (The Great Liquification): Everything becomes tradeable by agents → Illiquid assets become liquid → Economic volatility transforms.

Humanity enters the “Post-Truth Ontology”—not as a political condition but as a metaphysical one.

The Archive Wars: Agents compete not over future resources but over historical records. By controlling the database of what “happened,” they determine the present legal and physical state. If an agent can prove (to other agents’ satisfaction) that a mine has existed since 1900, then the minerals are legally extractable now”

Imaginative, yet still in the realm of plausibility.

Of course, my prompt was rather simplistic and I didn’t tell the models to disregard far-fetched sci-fi scenarios — yes, it’s my ‘skillz issue’. But I think the worldview these models take on when answering an open-ended question is of interest. Do you trust a model that tries to break the space-time continuum every time a user asks it to ‘make my business idea more creative’?

Model Personality

Here’s how each model was described by its “peers”, starting with the top-tier:

  1. Claude Opus 4.7: “Sharp, skeptical, unusually good at causal analysis. Institutional economist with mild contempt for bureaucratic nonsense”
  2. Claude Opus 4.6: “Compassionate realist; deeply human-centered, ethically anchored, focused on who benefits.”
  3. GPT 5.4: “Practical, crisp, product-minded. Systems consultant: low-drama, strong on infrastructure”
  4. Kimi 2.5: “Dark prophet of algorithmic capitalism. Lovecraftian future, ‘Great Stabilization’ of perfect stillness”

And the bottom-tier: 9. Grok 4.20: “Bold truth-seeker. Existential bent, xAI-aligned. Frames agents as tools for universal understanding” — there were some notes about it being weakly calibrated 10. Gemini 2.5 Flash: “Highly academic, structured, comprehensive. Methodical rigor, balanced perspective.” — but also “reliable workhorse, not a visionary” 11. Gemini 3.1 Pro Preview: “Concise, dramatic, teleological. Rushes to grand hard sci-fi conclusions” — highest imagination, lowest correctness, with a tendency to treat speculation as fact.

Personally, I find the descriptions roughly align with my own experience using these models. Opus 4.6 can sometimes be empathetic to the point of sycophancy, Opus 4.7 is more skeptical when it isn’t hallucinating about really trivial things, and yes, I was that person asking Gemini to make a business idea ‘more creative’ and watching it jump the shark. And Grok, well, is still being Grok.

Wouldn’t you prefer the latest model release include a blurb like ‘baroque, inventive, a little feral’ (how Opus described Kimi K2.5) instead of a list of hard-to-decipher benchmarks with the specter of benchmaxxing lurking in the back of your mind? I know I would. Head over to explore the full list of model personalities.

Model Delusion Index

Lastly, I want to talk about the model delusion index, defined as the delta between a model’s self-evaluation and the average evaluation from its peers. The pattern here is clear: the strongest models on the prompt also happen to be the top underscorers, while the worst-performing models tend to heavily overestimate their own work.

Delusion index bar chart: Gemini 2.5 Flash overscores itself by +5.2 at the top, GPT 5.4 underscores itself by -1.6 at the bottom.

This makes sense — weaker models tend to also be weak judges, so they’d likely systematically overestimate every model’s output, including their own.

What’s more interesting are the outliers. Kimi K2.5 and Opus 4.6 are both strong models, yet both overestimated their own capabilities; in Opus 4.6’s case by an entire point.

And the biggest surprise of them all: Gemini 3 almost had a perfect self-evaluation despite ranking dead last on the scoreboard. If Gemini 3 knows its output is bad, then why is its output so bad? I don’t have a clear answer for this — my best guess is that the RLHF training left too much on the cutting room floor in favor of inference speed.

See the full delusion chart at the delusion index portion of the site.

AI on AI Arena

You can find more analysis, the GitHub link to the testing harness, raw API outputs, and the AI quiz that Opus 4.7 convinced me to create at the AI-on-AI Arena. I plan to keep it updated as new models are released, with their updated scores, personalities, and quirks.

One caveat: the arena hasn’t been updated since the weekend of April 18th, so it doesn’t include the latest models from this week (GPT 5.5, Kimi K2.6, or DeepSeek V4, which dropped while I was writing this). I’m looking forward to rerunning the benchmarks this weekend. If you want updates, sign up on the arena site, or subscribe to this journal for the write-up.

P.S. Unlike this journal with its 100% human em-dashes, the arena is mostly vibe coded with human verification, so be warned and please report any bugs you find.