惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

cs.AI updates on arXiv.org
cs.AI updates on arXiv.org
罗磊的独立博客
B
Blog RSS Feed
C
Check Point Blog
Project Zero
Project Zero
D
Docker
CTFtime.org: upcoming CTF events
CTFtime.org: upcoming CTF events
P
Palo Alto Networks Blog
L
LINUX DO - 热门话题
Scott Helme
Scott Helme
NISL@THU
NISL@THU
L
LangChain Blog
C
Cisco Blogs
Engineering at Meta
Engineering at Meta
Know Your Adversary
Know Your Adversary
雷峰网
雷峰网
S
Schneier on Security
MyScale Blog
MyScale Blog
博客园_首页
博客园 - 三生石上(FineUI控件)
C
CERT Recently Published Vulnerability Notes
美团技术团队
V
Visual Studio Blog
T
The Exploit Database - CXSecurity.com
Recent Announcements
Recent Announcements
G
GRAHAM CLULEY
T
Tor Project blog
V
Vulnerabilities – Threatpost
U
Unit 42
Cyber Security Advisories - MS-ISAC
Cyber Security Advisories - MS-ISAC
Stack Overflow Blog
Stack Overflow Blog
P
Privacy International News Feed
Security Latest
Security Latest
W
WeLiveSecurity
aimingoo的专栏
aimingoo的专栏
Hugging Face - Blog
Hugging Face - Blog
Google Online Security Blog
Google Online Security Blog
V2EX - 技术
V2EX - 技术
The Last Watchdog
The Last Watchdog
博客园 - Franky
T
Tenable Blog
云风的 BLOG
云风的 BLOG
D
DataBreaches.Net
钛媒体:引领未来商业与生活新知
钛媒体:引领未来商业与生活新知
Latest news
Latest news
N
News and Events Feed by Topic
Cloudbric
Cloudbric
Schneier on Security
Schneier on Security
让小产品的独立变现更简单 - ezindie.com
让小产品的独立变现更简单 - ezindie.com
Blog — PlanetScale
Blog — PlanetScale

Business Tech News: Latest Updates on Innovations, Startups, and Market Trends | The HinduBusinessLine

Geo-engineering against climate change ZincGel vs Li-ion battery Why the energy sector isn’t AI-ready yet IT services giant TCS takes an AI-led avatar IIT-M revives forgotten route to industrial wastewater treatment IIT-Kanpur-incubated start-up develops unique battery technology Two faces of water Why the made-in-India ePlane is unique Moving satellite data at laser speed Longer-lasting zinc battery How simulation tech can ready robots for the real world DAE commissions world’s first nuclear heat-based copper-chlorine hydrogen plant DAE commissions world’s first nuclear heat-based copper-chlorine hydrogen plant Subterranean forest of fungi Using sound waves to bypass charge-based circuits AI aides to decode Indian law How the US funding cut impacts cancer research The time to deploy thorium is now The protein-peptide bonds that heal IIT-Kanpur hosts India’s first DORIS beacon How plants summon help Fishing out fake news using a deep-learning neural network IIT-Madras sets up testing tank for ships, submarines Dentistry’s prehistoric drill With AI, science is borderless How ‘spent’ graphite breathes new life into fuel cell Coal gas can yield clean hydrogen at $1.25 a kg Light, compact antennas IMD launches pilot weather forecast within 1 km radius in UP, national roll out in 2-3 years Nationwide ban soon on Paraquat herbicide over toxicity concerns, health risks ParvAI: ‘Windows to the soul’ and workplace safety Why agreeable AI is a liability in competitive markets Indian material for magnet making Using lasers to punch holes in cell walls When the grid becomes an all-knowing data system Micro-mining for critical rare earth minerals Half the capex, less carbon: The molten magic inside Tata Steel’s HIsarna bet Cosmic aid for miners Efficient brakes and EV range India contributes ₹745 crore to multi-country ITER Big budgets, slow science: BARC under-spends on R&D Artemis-2: Hurtling moon-ward on an epochal mission Power supply lessons for AI Why nuclear fusion is gaining funding Defence research stays underfunded Micro attacks on sewer lines Turning the ubiquitous optical fibre into a sensor The PRAGYA tokamak Mind-reading tech Carnot battery: Carbon dioxide as ideal ‘working fluid’ On a leash of light On a wing and an AI-powered tool How do ‘natural polypills’ work? AI tool for capturing and managing hospital records How sea microbes can protect agri fields Why India should choose to build not just powerful, but also governable AI Flaring and quaking Qualcomm has an Edge in India Soil testing of rhizosphere CMFRI achieves captive breeding of threatened mangrove clam No erasures RDI scheme could be operationalised this year IIT-M’s ramjet shell is an engineering marvel Sun-powered supercapacitor 10 years on, NALCO yet to start gallium extraction project Budget doubles allocation for nuclear research to ₹2,410 cr Underwater water Recent successes in science-led atmanirbharta Electric mobility may take wing in the not-too-distant future Eco-friendly semiconductors Twinning prayers and AI at mega temple festival Solar cells of efficiencies above 30% A lesson from Germany on infrastructure maintenance Fabled city in the high mountains Optimising bioreactor design Sensing UV-C in femtoseconds ISRO to kick off 2026 with launch of Earth Observation Satellite Thriving in extremes Indo-Lankan leg-up for S&T Using AI to better assess cyclone damage War on drug resistance goes undersea Big, bad business of junk food Rosatom’s mini variant of small modular reactor Clear thinking on pranayama Can GenAI be a responsible teaching assistant? Pharma PLI fetches ₹26,832 cr sales ‘Scripting’ ideal AI output Honeywell’s technology may bring biomass to the centre stage India-made human-like robot Scorched by 163-year drought NTT’s quantum leap into near sci-fi realm A reality check on AI’s negotiation skills Salinity-proof epoxy coating for marine installations Heat from small-scale solar units could accelerate India’s net-zero transition Cross-species transplantation is at a regulatory crossroads Nature, the ultimate climate warrior Breakthrough in desalination technology, using carbon ‘flowers’ Epidemiology-ML collab decodes India’s struggles with air quality
No exam is too hard for AI?
By N Nagaraj · 2026-03-23 · via Business Tech News: Latest Updates on Innovations, Startups, and Market Trends | The HinduBusinessLine

Someone, at some point, perhaps as a joke, decided to call it “Humanity’s last exam”, or HLE. By the time it was published in Nature in January, its designers had already announced a replacement. The replacement is updated continuously. It has to be.

It is a benchmark of 2,500 questions assembled by nearly 1,000 experts from 500 institutions across 50 countries, introduced by researchers at the Center for AI Safety and Scale AI. At launch, the best AI models scored under 10 per cent. The live leader board now shows 38.3 per cent. The trajectory, more than the absolute score, is the point.

The benchmark was designed in response to a failure in existing tools for measuring AI capability. Frontier models had already exceeded 90 per cent accuracy on MMLU (massive multitask language understanding), once considered a serious challenge. When all the best systems can clear it, the measuring instrument does not tell you much about the differences between them, or where they may be the following year.

HLE’s design logic was deliberately adversarial. Questions were submitted by experts across more than a hundred disciplines. Before entering the dataset, each had to defeat all current frontier models and pass two rounds of expert review. The result was a set of questions that, by construction, no existing AI could reliably answer.

The evaluation results confirmed the difficulty. At launch, GPT-4o scored 2.7 per cent, Claude 3.5 per cent, Sonnet 4.1 per cent, and OpenAI’s o1 8 per cent. DeepSeek-R1 reached 8.5 per cent. These numbers were low by design. The speed of subsequent progress was not fully anticipated. GPT-5 scored 25.3 per cent, Gemini 2.5 Pro reached 21.6 per cent. Scores have continued to climb; the live leader board now shows Gemini 3 Pro at 38.3 per cent.

Calibration errors across models ranged from 50 per cent to 89 per cent. Calibration measures whether a model’s stated confidence matches its accuracy. On HLE, models routinely expressed high confidence while being wrong. This is not a quirk of one system. It holds across architectures, suggesting a structural feature of current AI design.

Accuracy improved with more reasoning compute, but only to a point. Beyond roughly 16,000 output tokens, performance declined.

Dynamic testing

The paper reports an expert disagreement rate of 15.4 per cent, rising to 18 per cent in biology, chemistry and health. Nearly one in six questions could not be answered consistently even by specialists. The benchmark is harder than any existing AI can reliably handle, but also harder for any single human expert.

The authors draw one boundary clearly: High performance on HLE would not constitute evidence of artificial general intelligence. It would demonstrate expert-level performance on close-ended academic questions. That distinction is often lost in public commentary.

There is, however, a deeper problem. Every instrument built to measure AI capability has so far become a ceiling that models eventually reach. MMLU took years to saturate; HLE is showing pressure within months. The designers acknowledge this by announcing HLE-Rolling, a dynamically updated version intended to stay ahead of the models.

Once a benchmark is published, it becomes a target for developers and for the optimisation logic by which models are compared and sold. The instrument and the thing it measures cease to be independent. No static benchmark can escape this. The inability to build a stable yardstick suggests that capability is moving faster than the ability to define what is being measured.

Business decisions

For businesses, this has three consequences. Investment and procurement decisions across sectors are being made on capability claims built, directly or indirectly, on benchmark performance. If the benchmarks are structurally unstable, so are those claims.

The calibration finding compounds this. A model that cannot signal its own uncertainty is unsuitable for deployment in contexts where errors compound, including credit assessment, medical triage and document-intensive knowledge work. Confident wrongness is operationally worse than uncertain wrongness because it removes the incentive for a human check.

The pace of improvement also means that any organisation’s current understanding of AI capability has a short shelf-life. Planning and regulatory assumptions require revision cycles that most institutions are not designed to support.

The scores are rising fast enough; so it may not be possible to decide, in any meaningful way, which model is leading. At the margins of a benchmark this hard, differences between the best systems will fall within statistical noise. The yardstick will not so much have been beaten as dissolved.

More Like This

Published on March 23, 2026