惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

AI
AI
Scott Helme
Scott Helme
W
WeLiveSecurity
N
News | PayPal Newsroom
G
GRAHAM CLULEY
SecWiki News
SecWiki News
V2EX - 技术
V2EX - 技术
Security Latest
Security Latest
H
Heimdal Security Blog
L
LINUX DO - 最新话题
Application and Cybersecurity Blog
Application and Cybersecurity Blog
C
Cyber Attacks, Cyber Crime and Cyber Security
P
Palo Alto Networks Blog
Simon Willison's Weblog
Simon Willison's Weblog
PCI Perspectives
PCI Perspectives
钛媒体:引领未来商业与生活新知
钛媒体:引领未来商业与生活新知
Hugging Face - Blog
Hugging Face - Blog
博客园_首页
Spread Privacy
Spread Privacy
T
Troy Hunt's Blog
V
Vulnerabilities – Threatpost
罗磊的独立博客
C
CXSECURITY Database RSS Feed - CXSecurity.com
宝玉的分享
宝玉的分享
T
The Blog of Author Tim Ferriss
阮一峰的网络日志
阮一峰的网络日志
Martin Fowler
Martin Fowler
The Hacker News
The Hacker News
C
Cisco Blogs
T
Tor Project blog
cs.CV updates on arXiv.org
cs.CV updates on arXiv.org
The Register - Security
The Register - Security
Cyber Security Advisories - MS-ISAC
Cyber Security Advisories - MS-ISAC
P
Privacy & Cybersecurity Law Blog
奇客Solidot–传递最新科技情报
奇客Solidot–传递最新科技情报
S
Security Affairs
T
Tenable Blog
V
Visual Studio Blog
C
Check Point Blog
Hacker News - Newest:
Hacker News - Newest: "LLM"
CTFtime.org: upcoming CTF events
CTFtime.org: upcoming CTF events
美团技术团队
月光博客
月光博客
J
Java Code Geeks
量子位
Vercel News
Vercel News
I
Intezer
博客园 - 聂微东
Know Your Adversary
Know Your Adversary
aimingoo的专栏
aimingoo的专栏

Machines – Silicon Republic

Samsung’s new robotics division appoints former Boston Dynamics lead Moonshot AI pauses subscriptions after users swarm to test Kimi K3 Alibaba unveils Qwen3.8, 'second only to Fable 5' Moonshot unveils Kimi K3, largest open-weight AI model yet Mira Murati’s AI start-up unveils customisable model Inkling Ireland looking ahead with launch of Quantum 2030 Implementation Plan Economists and tech leaders sign statement warning of AI threats How to ensure digitisation in healthcare doesn’t flatline Meta enters pay-to-use AI market with Muse Spark 1.1 Altman says new GPT-5.6 model 54pc more token-efficient Mistral expands physical AI offering with first robotics launch Apple and Broadcom renew chip deal, which will run to 2031 Infineon opens new €5bn Smart Power Fab chip factory in Germany Finnish quantum company IQM makes history with Nasdaq debut Ireland secures €10m to launch AI Factory Antenna Anthropic launches Claude Science app for researchers and scientists Google Cloud Marketplace to offer LQMs from SandboxAQ IBM unveils tech capable of producing chips smaller than one nanometre Dublin's TensorX to partner with Solstice on sovereign European AI If AI robots can be tricked into ‘going rogue’, what are the implications? Tencent tests new AI agent Xiaowei on WeChat 3,900 Waymo robotaxis recalled after new software issue New Irish bill to supervise EU AI Act gets greenlit Horizon Quantum to build second quantum computer in Dublin Anthropic rolls out ‘Mythos-like’ AI model Claude Fable 5 EU AI Act – the high-risk classification guidelines explained Will AI ‘digital twins’ transforming heart care work for women? Past the ‘wow phase’ of robotics, delivery and safety are paramount AI changing jobs faster than companies can keep up with, finds report Kerry’s RDI Hub opens AI collaboration with Luxembourg IQM raises PIPE to $146m with Finnish pension fund backing Huawei proposes new path for chips as Moore’s Law runs out of road France bets fresh €1bn on quantum as global race intensifies US pumps $2bn into quantum computing via CHIPS Act As AI meets science, what is in store for the future of research? Dublin’s Ubotica teams with Novi for real-time orbital data analysis Irish quantum start-up Equal1 unveils RacQ data centre computer Waymo trouble: 3,800 robotaxis recalled after software glitch OpenAI launching security AI initiative to compete with Claude Mythos Opinion: Europe can’t afford to sit on the agentic commerce sidelines Moonshot AI valued at $20bn after $2bn raise for Kimi creator ServiceNow wants to be ‘AI agent of agents’ with Otto platform and AI tools Bloomberg: China pauses AV permits after Baidu disruption Milestone reached in Celtic Interconnector project linking France and Ireland Semiconductors core to Tyndall's five-year strategy China's DeepSeek unveils long-awaited V4 AI model Can you rely on AI chatbots for medical advice? Merz, Siemens call for easing of EU regulations on industrial AI Are electric vehicles about to take off for good? OpenAI to rival Google’s AlphaFold with new AI model for life sciences research Are we ready to place lab experiments in non-human hands? Nvidia unveils open-source quantum AI model Ising Stanford: China ‘effectively’ closes AI model performance gap to US Opinion: The future of insurance is AI, so why the hesitation? Anthropic reportedly mulls designing own chips amid shortage Equal1 partners with Q-Ctrl for quantum data centre deployment Meta’s Superintelligence Labs debuts first product Muse Spark Agentic commerce and purchase disputes: Did you mean to buy that? Anthropic, Google, Broadcom announce 3.5GW TPU deal Microsoft releases foundational AI models targeting enterprises What issues arise when code has the ability to write and review itself? France buys supercomputer maker Bull in tech sovereignty push Anthropic accidentally leaks Claude Code source in npm slip The deep-tech founder using AI to address immunology challenges New German battery recycling plant salvages lithium and graphite Irish AI start-up Jentic joins OpenClaw push with Jentic Mini Amazon acquires humanoid robot start-up Fauna Robotics Policy as code: Embedding compliance in AI adoption Is AGI really here as Nvidia’s Jensen Huang claims? Alibaba International launching agentic AI for enterprise Anthropic takes on OpenClaw with new Claude Code text feature
‘AI scientists’ are improving, but what are the fundamental limits?
silicon · 2026-05-27 · via Machines – Silicon Republic

Karin Verspoor of RMIT University explores how AI is impacting research in STEM.

Many of the most exciting discoveries in science involve highly specialised knowledge and making connections between far-flung facts. Scientists must combine deep analysis with broad reasoning strategies.

As in many information-rich tasks, researchers are looking to artificial intelligence (AI) systems to speed up their work. AI tools may be able to support key steps such as generating ideas, reviewing existing work and analysing data.

The latest systems use large language models (LLMs) to allow scientists to interact naturally and directly with the vast body of knowledge captured in words in the scientific literature.

But as two new systems described in papers just published in Nature show, when it comes to science, language alone can only go so far.

What AI is doing to science

A number of organisations, such as Sakana AI, are trying to automate the entire scientific process. To date, these efforts have largely focused on computer science, where ‘experiments’ mainly involved designing and writing code.

However, the Agents4Science conference organised at Stanford last October showcased a broader range of AI-generated papers. They covered topics from mechanical engineering and protein design to a system called BadScientist which deliberately produced “convincing but unsound” research.

I have previously raised concerns about the impacts of AI scientists on the scientific ecosystem. Recent work validates these concerns, showing increased quantity but lower quality of both papers and peer reviews, identifying fabricated references in published works, finding fabricated and misleading images, and more.

What scientists are doing with AI

AI systems clearly can’t be trusted to conduct the full process of science on their own. But how about using AI to help scientists get more done more quickly?

This is the intent of the two new systems described in Nature: Robin, made by non-profit Future House, and Co-Scientist, from Google DeepMind.

Both systems aim to accelerate scientific discovery, working in collaboration with a scientist. Both are also ‘multi-agent’ AI systems, meaning they are built as a collection of specialised agents each targeting specific steps of the scientific discovery process, coordinated by a ‘supervisor’ agent.

The agents that comprise Co-Scientist aim to mirror abstract cognitive tasks, such as a ‘reflection agent’ that acts as a critical scientific peer reviewer assessing the quality of a hypothesis. ‘Ranking agents’ debate research hypotheses in ‘tournaments’, using multiple interacting LLMs to simulate a discussion about the relative merits of two hypotheses.

Robin’s agents, on the other hand, are more tuned to specific tasks relevant to drug repurposing, aiming to identify new drugs for a given disease. One agent focuses on selecting experimental tests, while another analyses complex biomedical data.

How do the results stack up?

Co-Scientist can assess the quality of its generated proposals, using a method called the Elo rating which is best known for ranking chess players. Co-Scientist’s self-ratings of the novelty and impact of its outputs align quite well with the preferences of human experts and judgements by other LLM systems.

In a drug repurposing experiment, Co-Scientist selected 30 drug candidates as promising treatments for a kind of cancer called acute myeloid leukemia. Expert (human) oncologists refined the list, and five drugs were tested in the lab. Of these, three showed some positive results and one seemed to show particular promise.

Other experiments showed the potential of Co-Scientist to explore combinations of multiple drugs.

Notably, the predictions of Co-Scientist were not compared with the plethora of targeted computational and machine learning methods for drug repurposing that have been developed over decades of computational biology research. This means we don’t know whether the new general-purpose tool outperforms more specific AI approaches.

Both systems stop short of validating their hypotheses directly, which would involve real physical experiments. Both also rely heavily on human input to define the key scientific question, sense-check predictions and prioritise predictions for further investigation.

Co-Scientist focuses primarily on generating hypotheses through elaborate reasoning agents, leaving validation and interpretation to subsequent steps. Robin also uses an agent to analyse data produced from real-world experiments.

Robin was used to propose 30 drug candidates for a condition called dry age-related macular degeneration. The top five were selected for testing.

Robin also made proposals for the experiments, with several suggestions overridden by the human scientists. Through several rounds of brainstorming and analysis, two drugs were identified as promising.

Testing of Robin’s individual agents showed those that dug through earlier research were better at the task than general-purpose LLMs. The analytical agent did less well on questions about statistics and bioinformatics, and relied heavily on human-supplied prompts.

The limits of language alone

AI can help scientists to navigate the vast amount of documented knowledge humans have acquired over the millennia. Use of computation to find patterns in large datasets, to integrate dispersed information, and to drive new discoveries from existing literature has already contributed to scientific progress for decades.

New models such as Robin and Co-Scientist represent a shift towards working directly in the realm of the language of science, rather than the realm of raw data. This allows more natural collaborations between scientist and machine, through language-based ‘discussions’.

However, more natural doesn’t necessarily mean more effective. Language-based communication can be imprecise and ambiguous, where science must be specific.

Models that combine the best of these worlds are on the horizon. These aim to link structured quantitative data to the concepts and relationships that describe the core facts beneath it.

Such models ground scientific reasoning in the structure of knowledge. They allow scientific evidence ranging from genomic sequences and protein structures to cellular imaging to be connected.

Words are how science is communicated. AI tools that facilitate making sense of the information that is hidden in all of those words are surely valuable. But the complexity of the natural world means that AI (co-) scientists will only be truly effective when they can go beyond connecting words together, to modelling the full complexity of the systems those words describe.

The Conversation

By Karin Verspoor

Karin Verspoor is dean of the School of Computing Technologies at RMIT University. She works at the intersection of science and technology, applying artificial intelligence methods to analysis and interpretation of biological and clinical data, focusing particularly on natural language processing of unstructured text data. She is a fellow of the Australian Academy of Technological Sciences and Engineering and the Australasian Institute of Digital Health. She is also the co-founder and Victoria node lead of the Australian Alliance for Artificial Intelligence in Healthcare.

Don’t miss out on the knowledge you need to succeed. Sign up for the Daily Brief, Silicon Republic’s digest of need-to-know sci-tech news.