惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

博客园 - Franky
J
Java Code Geeks
腾讯CDC
奇客Solidot–传递最新科技情报
奇客Solidot–传递最新科技情报
Jina AI
Jina AI
博客园 - 司徒正美
Stack Overflow Blog
Stack Overflow Blog
美团技术团队
L
LangChain Blog
WordPress大学
WordPress大学
A
About on SuperTechFans
Martin Fowler
Martin Fowler
月光博客
月光博客
Y
Y Combinator Blog
U
Unit 42
D
Docker
Recent Announcements
Recent Announcements
Hugging Face - Blog
Hugging Face - Blog
B
Blog
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻
G
Google Developers Blog
Last Week in AI
Last Week in AI
T
The Blog of Author Tim Ferriss
让小产品的独立变现更简单 - ezindie.com
让小产品的独立变现更简单 - ezindie.com

Hacker News - Newest: "AI"

AI can't read an investor deck AI as an attorney? Student uses ChatGPT, Gemini to sue UW over alleged racial discrimination Hacking MCP Servers in AI Systems – The Rug Pull: Tool Changes After Approval GitHub - MeepCastana/KubeezCut: Free Web based video editor Can AI judge journalism? A Thiel-backed startup says yes, even if it risks chilling whistleblowers Coming soon: 10 Things That Matter in AI Right Now DARPA built an AI to fact-check enemy weapons claims What explains heterogeneity in AI adoption? When AI Meets Muscle: Context-Aware Electrical Stimulation Promises a New Way to Guide Human Movements - Department of Computer Science AI Changed How We Build. It Did Not Change What Matters. Linux rules on using AI-generated code - Copilot is OK, but humans must take 'full responsibility for the… Meta spins up AI version of Mark Zuckerberg to engage with employees Code Mode: Let Your AI Write Programs, Not Just Call Tools | TanStack Blog GitHub - Delavalom/graft: Go framework for building AI agents. Type-safe tools, multi-provider (OpenAI, Anthropic, Gemini, Bedrock), zero vendor SDKs. India's TCS tops estimates, says new AI models did not dent services demand Gen Z's fading AI hype Strong feeling: we are in a folded AI reality GitHub - machinarii/total-recall-catalog: A reference catalog of latest knowledge retrieval, memory & RAG systems GitHub - mensfeld/code-on-incus: Give each AI agent its own isolated machine with root, Docker, and systemd. Active defense detects and stops threats automatically.. Quantization, LoRA, and the 8% Problem: Benchmarking Local LLMs for Production AI Iran war: We spoke to the man making Lego-style AI videos that experts say are powerful propaganda Powell, Bessent discussed Anthropic's Mythos AI cyber threat with major U.S. banks GitHub - immartian/bellamem: Persistent belief-graph memory for AI agents. Retrieves decisive context by importance — not recency, not RAG, not /compact. recursive-mode: The Repo-Native Operating System for AI Engineering After the attack on Sam Altman's home, will AI CEO's go on the offensive? The biggest advance in AI since the LLM Opus 4.6 vs GPT 5.4 One Prompt Unity World Generation Test “AI polls” are fake polls Client Challenge Can AI be a 'child of God'? Inside Anthropic's meeting with Christian leaders
Why AI can’t be trusted to write scientific reviews
Sarkar, Rupa · 2026-05-26 · via Hacker News - Newest: "AI"

Artificial-intelligence tools are being touted as a means to conduct rapid reviews of scientific literature. At the London-based publisher the Cochrane Collaboration, where I became editor-in-chief in March, we specialize in health-related systematic reviews: the highest-quality syntheses of all the available research. We are testing ways to use AI to increase our reviews’ efficiency and scale. But, in our experience, the current tools are far from ready for mainstream adoption, and the assumption that machines can replace humans on all methodological tasks is flawed.

The stakes are high. Systematic reviews and other types of evidence synthesis inform clinical practice, public-health guidance and policy decisions that affect entire populations. Errors could give false hope to patients or lead health systems to waste money on ineffective or unsafe interventions.

Current AI models typically replicate the step-by-step processes by which people conduct systematic reviews: identifying suitable studies from disparate sources; extracting relevant data for analysis; and, finally, writing up the report. The idea is to replace the work of humans.

But conducting systematic reviews is not a purely computational task. Human specialists are needed to define meaningful review questions, evaluate relevance, interpret results and understand clinical or policy implications. Context and subjective nuance are seldom well-represented in AI models’ training data, and the models’ tendency to hallucinate — that is, to fabricate information — means that their outputs need to be verified by human experts.

Efforts at Cochrane show the limitations of using AI in place of people. We’ve been evaluating tools that support study screening and data extraction. These are time-consuming processes to conduct manually, particularly when primary data are not readily accessible and must be drawn from multiple sources or inferred from published results.

But we’ve found that most of the tools available were developed by private companies. This is problematic for reviews that evaluate drugs and medical devices, because these need to be independent of industry. What’s more, few AI models are open source, with most relying on opaque, proprietary ‘black box’ processes. This means there’s no way to examine whether a tool might disproportionately include trials with results favourable to one drug company.

And, on the practical side, tools in the current generation require long training periods for both the AI and the human operator before they yield reliable results. So far, we’ve found that, for each review, the whole process takes longer than doing the work manually.

In my view, to realize the full potential of AI, it’s crucial that tool developers, authors and evidence users move away from using it to generate individual reviews. Instead of mimicking human processes, developers should start building systems that allow humans and AI to work together effectively to assess studies.