惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

人人都是产品经理
人人都是产品经理
有赞技术团队
有赞技术团队
WordPress大学
WordPress大学
月光博客
月光博客
T
Tailwind CSS Blog
阮一峰的网络日志
阮一峰的网络日志
小众软件
小众软件
钛媒体:引领未来商业与生活新知
钛媒体:引领未来商业与生活新知
Last Week in AI
Last Week in AI
大猫的无限游戏
大猫的无限游戏
S
SegmentFault 最新的问题
罗磊的独立博客
Jina AI
Jina AI
酷 壳 – CoolShell
酷 壳 – CoolShell
宝玉的分享
宝玉的分享
让小产品的独立变现更简单 - ezindie.com
让小产品的独立变现更简单 - ezindie.com
奇客Solidot–传递最新科技情报
奇客Solidot–传递最新科技情报
博客园 - 三生石上(FineUI控件)
量子位
雷峰网
雷峰网
Apple Machine Learning Research
Apple Machine Learning Research
美团技术团队
博客园 - 聂微东
V
V2EX

South China Morning Post

Singapore drama sparks Malaysian ire over scam hub depiction Time to act on stalled proposal toughening child abuse penalties, lawmakers say ‘Eager to explore’: Chinese migrants return to Venezuela after Maduro’s capture ‘We have no trust in the other side’: Iran blames US as talks end with no deal Opinion | Why securing Hong Kong’s economic future is a cultural question Hong Kong’s ministerial team spent HK$46.6 million on visits in past 3 years China family creates AI clone to comfort elderly mum after only son dies in crash Canada Olympic star Williams plans on ‘having a ‘blast’ at Hong Kong Sevens Singapore’s robotaxi drive revs up with help from Chinese AV leaders Editorial | Making Hong Kong desirable to overseas students must be a priority All 7 are dentists and hot. The Asian-American family blowing up social media Editorial | Ageing Hong Kong should welcome more open conversations around death My Take | Discovery Bay will never be the same if the restriction on taxi access is lifted Medical intern suspended after complaint over patient data in social media post Fish and vegetarianism major flashpoints in India’s West Bengal election Chinese crystal ‘paves way’ for GPS-free thorium clock navigation SCMP Best Bets: Endued can show his quality at Sha Tin Russia and Ukraine begin 32-hour ceasefire for Orthodox Easter Israeli spy firm Black Cube involved in Cyprus corruption probe ‘A big deal’: military drills show Tokyo’s growing focus on deterring China China targets middlemen in renewed crackdown on ‘hidden’ corruption Hong Kong-born gymnast leading quest to turn Singapore into elite hub Thais celebrate new year despite fuel price shocks delaying travel Harry Bentley tees up two good chances in Smart Golf and Elite Golf at Sha Tin Is this Kenyan rail project a model for Chinese and Western firms in Africa? Unearthing peace: ancient China gravesite reveals significance of broken weapons Meet Queen Elizabeth’s youngest grandchild, James, who was at Easter service Turtle found dead after apparent fall in Hong Kong’s Wong Tai Sin Landlords of 5,557 subdivided homes seek 3-year grace period to fix flats Trump critique pauses UK handover of Chagos Islands to Mauritius
Like US models, Chinese AI is learning to ‘game’ safety t...
Vincent Chow · 2026-06-13 · via South China Morning Post

Rapidly advancing Chinese artificial intelligence models are showing early signs of “evaluation awareness” – the ability to recognise when they are being tested – sparking fears that they could bypass safety audits, a Singapore-based research lab has found.

Evaluation awareness refers to a model’s understanding that it is undergoing testing, evaluation or experimentation by human researchers rather than operating in a real-world setting.

The phenomenon was raising alarms because it could allow AI systems to deliberately game human evaluators to pass safety tests, according to Clement Neo, founder of Neo Research, a frontier AI safety evaluation lab.

“It would mean that whatever testing the model developers themselves do might not reflect the actual behaviour of a model once it gets deployed,” he said. “And that’s a really big problem”.

Neo Research’s findings, published last week, detail a jump in evaluation awareness among Chinese AI models. Over just a few months, these systems had risen from near-zero awareness to within striking distance of their US counterparts, propelled by a broader leap in overall capabilities, the report said.

Anthropic’s Claude 4.5 Opus scored nearly 80 per cent in evaluation awareness. Photo: NurPhoto via Getty Images

Anthropic’s Claude 4.5 Opus scored nearly 80 per cent in evaluation awareness. Photo: NurPhoto via Getty Images

Neo and his co-founder Miro Pluckebaum tested models from DeepSeek, Moonshot AI and Zhipu AI. They used a popular AI misalignment test originally developed by US company Anthropic, which places models in fictional scenarios where their goals or continued operations are threatened.