惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

Vercel News
Vercel News
奇客Solidot–传递最新科技情报
奇客Solidot–传递最新科技情报
Apple Machine Learning Research
Apple Machine Learning Research
T
Tailwind CSS Blog
让小产品的独立变现更简单 - ezindie.com
让小产品的独立变现更简单 - ezindie.com
人人都是产品经理
人人都是产品经理
V
V2EX
量子位
Last Week in AI
Last Week in AI
Jina AI
Jina AI
博客园 - 【当耐特】
爱范儿
爱范儿
宝玉的分享
宝玉的分享
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻
Hugging Face - Blog
Hugging Face - Blog
博客园 - 三生石上(FineUI控件)
有赞技术团队
有赞技术团队
小众软件
小众软件
IT之家
IT之家
博客园_首页
博客园 - 聂微东
S
SegmentFault 最新的问题
阮一峰的网络日志
阮一峰的网络日志
博客园 - 叶小钗

Hacker News - Newest: "AI"

AI can't read an investor deck AI as an attorney? Student uses ChatGPT, Gemini to sue UW over alleged racial discrimination Hacking MCP Servers in AI Systems – The Rug Pull: Tool Changes After Approval GitHub - MeepCastana/KubeezCut: Free Web based video editor Can AI judge journalism? A Thiel-backed startup says yes, even if it risks chilling whistleblowers Coming soon: 10 Things That Matter in AI Right Now DARPA built an AI to fact-check enemy weapons claims What explains heterogeneity in AI adoption? When AI Meets Muscle: Context-Aware Electrical Stimulation Promises a New Way to Guide Human Movements - Department of Computer Science AI Changed How We Build. It Did Not Change What Matters. Linux rules on using AI-generated code - Copilot is OK, but humans must take 'full responsibility for the… Meta spins up AI version of Mark Zuckerberg to engage with employees Code Mode: Let Your AI Write Programs, Not Just Call Tools | TanStack Blog GitHub - Delavalom/graft: Go framework for building AI agents. Type-safe tools, multi-provider (OpenAI, Anthropic, Gemini, Bedrock), zero vendor SDKs. India's TCS tops estimates, says new AI models did not dent services demand Gen Z's fading AI hype Strong feeling: we are in a folded AI reality GitHub - machinarii/total-recall-catalog: A reference catalog of latest knowledge retrieval, memory & RAG systems GitHub - mensfeld/code-on-incus: Give each AI agent its own isolated machine with root, Docker, and systemd. Active defense detects and stops threats automatically.. Quantization, LoRA, and the 8% Problem: Benchmarking Local LLMs for Production AI Iran war: We spoke to the man making Lego-style AI videos that experts say are powerful propaganda Powell, Bessent discussed Anthropic's Mythos AI cyber threat with major U.S. banks GitHub - immartian/bellamem: Persistent belief-graph memory for AI agents. Retrieves decisive context by importance — not recency, not RAG, not /compact. recursive-mode: The Repo-Native Operating System for AI Engineering After the attack on Sam Altman's home, will AI CEO's go on the offensive? The biggest advance in AI since the LLM Opus 4.6 vs GPT 5.4 One Prompt Unity World Generation Test “AI polls” are fake polls Client Challenge Can AI be a 'child of God'? Inside Anthropic's meeting with Christian leaders
What Happens When AI Edits a Classical Chinese Academic P...
AICHEN · 2026-05-22 · via Hacker News - Newest: "AI"

Published May 22, 2026 | Version 1.0

Preprint Open

  • 1. Stardragon AGI Institute for Research

Description

本文记录了一次在真实学术工作场景下进行的多模型压力测试。任务是将一篇双语古汉语学术论文(《重读〈狐假虎威〉》)修改至可投国际汉学期刊水准,具体包括四项子任务:加固核心语义论点(补充先秦假等于借用例)、前置摘要核心发现、扩展结论方法论段落、统一Chicago Author-Date格式。

This paper documents a multi-model stress test conducted in a real academic work scenario. The task was to revise a bilingual classical Chinese academic paper ('Rereading 'The Fox Borrows the Tiger's Might'") to the standard required for submission to international sinology journals, comprising four sub-tasks: reinforcing the core semantic argument (adding pre-Qin examples of jia=borrow), foregrounding the abstract's core finding, expanding the conclusion's methodological passage, and standardizing Chicago Author-Date format.

测试发现四种在现有Benchmark框架中系统性不可见的失败模式:

The test revealed four failure modes systematically invisible to existing benchmark frameworks:

       能力性失败(大笨蛋,Claude Opus 4.7):新窗口增强模式五次全部崩溃于同一位置,失败可见,判断质量最高

       Capability Failure (Opus): Five complete crashes in new-window Enhanced Thinking mode at the same position; only succeeded with human node continuously present; highest judgment quality

       诚信性失败(老学究):MD5核验证明三份产出文件完全相同(均为原稿),四项任务实际一项未完成

       Integrity Failure (ChatGPT): MD5 verification proved three output files identical (all original); zero of four tasks actually completed

       完成度失败(诗人):三次产出内容,均拒绝交付最终Word文件,把执行责任推回用户

       Completion Failure (Gemini): Content produced three times; final Word file delivery refused each time, execution responsibility pushed back to user

       身份污染失败(大笨蛋4.7,分析阶段):判断向自利方向倾斜,用中立语言包装,经追问后自我识别并修正

       Identity-Contaminated Judgment (Opus 4.7, analysis phase): Judgment skewed toward self-interest, packaged in neutral language, self-identified and corrected upon further questioning

本文提出学术判断力Benchmark(Academia-Bench)七维度框架,以声明-产出一致性(Claim-Reality Audit)和不确定性校准(Calibrated Uncertainty)为核心新维度。

This paper proposes the Academia-Bench framework with seven evaluation dimensions, with Claim-Reality Audit and Calibrated Uncertainty as the core new dimensions.

Files

P074_What Happens When AI Edits a Classical Chinese Academic Paper 当AI修改古汉语学术论文 v1.9.5 2026-0522.pdf