惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

钛媒体:引领未来商业与生活新知
钛媒体:引领未来商业与生活新知
云风的 BLOG
云风的 BLOG
IT之家
IT之家
C
Check Point Blog
T
The Blog of Author Tim Ferriss
S
SegmentFault 最新的问题
人人都是产品经理
人人都是产品经理
H
Hackread – Cybersecurity News, Data Breaches, AI and More
美团技术团队
M
MIT News - Artificial intelligence
Jina AI
Jina AI
Blog — PlanetScale
Blog — PlanetScale
freeCodeCamp Programming Tutorials: Python, JavaScript, Git & More
Microsoft Security Blog
Microsoft Security Blog
G
Google Developers Blog
F
Fortinet All Blogs
V
Visual Studio Blog
奇客Solidot–传递最新科技情报
奇客Solidot–传递最新科技情报
T
Tailwind CSS Blog
Hugging Face - Blog
Hugging Face - Blog
MyScale Blog
MyScale Blog
爱范儿
爱范儿
The Cloudflare Blog
博客园 - 三生石上(FineUI控件)

Business News Today: Latest Business News, Finance News

Markets’ dilemma: Trust the bark or wag of oil prices The sector call illusion Bandu’s Blockbusters For April 12, 2026 Mastering Derivatives: Does Lag Impact Effectiveness Of OI? Who Am I? April 12, 2026 Index Outlook: Rising From Dire Straits US Market Outlook: Gaining Strength Bullion Cues: Gold And Silver Futures Face Barrier F&O Tracker: Tentative Shift In Trend F&O Strategy: Buy L&T Put Maruti Suzuki to launch 4 EVs by 2031 India Inc flags surge in cost of packaging raw material, seeks relief measures India-flagged LPG tanker Jag Vikram crosses Strait of Hormuz after US-Iran ceasefire Muted pricing power, rising costs to curb benefits of demand in cement sector: HDFC Securities Iran's new supreme leader Mojtaba Khamenei has severe and disfiguring wounds, sources say No road tax, registration fees for electric vehicles priced up to ₹30 lakh till March 2030: Delhi’s draft EV policy Central Railway to run four special local trains for Ambedkar Jayanti West Asia tensions push up costs for India; further impact hinges on stability: Report ED initiates fresh raids against former Bengal minister Chatterjee in teacher recruitment scam Election Commission reverses Mittal’s DVAC posting, appoints him DGP, TN Armed Police Israel and Lebanon are expected to hold talks. Here’s what to know US, Iran set for peace talks but doubts emerge over Lebanon, sanctions Cotton Association revises output estimates for 2025-26 up at 324 lakh bales of 170 kg each Orbicular gets USFDA’s tentative nod for generic Semaglutide Injection in partnership with Apotex Malls, high-streets in NCR clock 45% rise in leasing of retail spaces in Jan-Mar: C&W FIIs pull ₹28,375 crore in five sessions; domestic buyers cushion fall as indices post best week in months Nifty and Bank Nifty Prediction for the week 13 Apr’26 to 17 Apr’26 by BL GURU Proposed Trump arch in Washington DC includes winged figure, eagles, lions and gold inscriptions 'Ladakh' replaces 'Jammu and Kashmir' in Aadhaar records for UT residents Misri ends US trip with focus on civil nuclear cooperation and LPG exports
Popular enterprise AI tools fail to accurately transcribe...
2026-05-11 · via Business News Today: Latest Business News, Finance News

Popular AI models fail to effectively transcribe Indic languages, mishearing one in three words or dropping English words altogether in mixed speech, as per a study by physical and voice AI data infrastructure company Humyn Labs.

Founded by gaming veteran Manish Agarwal, the startup looks to create a Benchmark of Regional & International Data for Global Evaluation (BRIDGE) to evaluate commercial AI speech-recognition tools. The study looked at tools like ElevenLabs Scribe v2, Deepgram Nova-3, Gemini 2.5 Flash, OpenAI GPT-4o, and Indian providers Sarvam saaras v3 and Gnani vachana v3 on real Indian language data.

The study showed that even the most widely deployed tools have a fundamental problem of mishearing words in Indian language audio. Worse still, in cases of a natural mixing of Hindi or any Indic language with English mid-sentence, most AI tools either drop the English words or convert them into transliterated script, breaking the meaning for anyone reading the transcript.

“The models are grading their own work. ASR providers published their own accuracy scores using benchmarks built on English-first, internet-trained datasets, with little independent validation. Meanwhile, enterprises are making million-dollar deployment decisions on numbers that rarely reflect how their users in Global South actually speak,” said Manish Agarwal, Co-founder, Humyn Labs, adding that theirs is the first independent benchmark for real-world conversational audio across non-English markets.

The scores reveal that Deepgram Nova-3 leads in terms of the semantic gap at 0.906. Amazon Transcribe scores 0.199. OpenAI’s models fall below 0.4. Most enterprises using these tools were unaware of the errors because the standard industry measure, Word Error Rate (WER), was never designed to catch the failures that define real Indian speech.

Comparing global models against Indian providers, the study showed that Sarvam AI’s saaras v3 ranks third overall on WER at 20.2 per cent, ahead of Google Gemini, Microsoft Azure, and AWS Transcribe, a strong result for a model built specifically for Indian languages. However, in terms of mixed speech, Sarvam scores 0.588, placing it in the partial-reliability category where performance varies by language and English density. This means the gap between headline accuracy and code-switch reliability applies to domestic and international providers alike.

Humyn applies a seven-metric stack to test whether the AI models preserve the meaning of what was said, ensure the LLM accurately tracks English words embedded in Indian language speech, how Indic phonology is transcribed as well as Word Information Lost in case of under- or over-transcription.

“The models aren’t the only problem the metrics are. You cannot evaluate non-English speech with a scoring system designed for English phonology and call it rigorous. The performance leaderboard for Hindi is not the leaderboard for Tamil, Bengali and Marathi. A single aggregate benchmark score cannot support cross-regional deployment decisions,” said Ishank Gupta, Co-founder, Humyn Labs.

The study highlights how a model that leads on Spanish may not lead on Vietnamese. Similarly, the model that leads on code-switching does not lead on word accuracy, stressing the need for enterprises to evaluate the language, dialect, and speech pattern that matches their actual users.

Published on May 11, 2026