惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

L
LangChain Blog
酷 壳 – CoolShell
酷 壳 – CoolShell
雷峰网
雷峰网
量子位
V
V2EX
S
SegmentFault 最新的问题
月光博客
月光博客
博客园 - 【当耐特】
Hugging Face - Blog
Hugging Face - Blog
V
Visual Studio Blog
大猫的无限游戏
大猫的无限游戏
T
Tailwind CSS Blog
博客园_首页
博客园 - Franky
钛媒体:引领未来商业与生活新知
钛媒体:引领未来商业与生活新知
美团技术团队
Y
Y Combinator Blog
The Cloudflare Blog
C
Check Point Blog
Cyber Security Advisories - MS-ISAC
Cyber Security Advisories - MS-ISAC
腾讯CDC
B
Blog
Stack Overflow Blog
Stack Overflow Blog
P
Proofpoint News Feed

Latest InfoTech News, IT, Information Technology News | The HinduBusinessLine

Google debuts standalone Gemini App for Apple’s MacOS India’s electronics imports cross $116 billion in FY26, exports trail Labour Ministry to look into POSH compliance by IT services firms, says employees union Is TCS harassment case tip of the iceberg? Entry-level smartphones get costlier as memory shortage persists Indians most nervous about AI despite highest skill penetration EdgeCortix secures Axiro, MPower investment to accelerate AI chip innovation Infosys partners Carlos Alcaraz as global brand ambassador Wipro buys select Alpha Net Group contracts for $70.8 Mn AMS expands Pune GCC, strengthens India’s role in global talent operations Memory chip crunch and Iran war lead phone market decline, IDC says UST, Evaaya jointly launch UST Nimbus to help empower GCCs with new capabilities No layoffs, says Zoho: 300 mentioned in social media post were interns Nvidia’s New AI models spark rally in quantum computing stocks OpenAI unveils GPT-5.4-Cyber a week after rival's announcement of AI model IMF urges nations to stay at frontier of mounting AI risks How your CCTV becomes a hacker’s spy Vehant Technologies eyes 20% topline from export in 3 years Cabinet Secretary emphasizes AI development and civil-military cooperation Amazon to acquire Globalstar for $11.57 billion to boost satellite internet Kaar Tech eyes data analytics acquisition as it positions itself as an AI-led enterprise OS enabler India’s quantum mission to complete 2,000 km network by 2027 Andhra Pradesh launches India’s first quantum reference facility in Amaravati Wegovy-maker Novo Nordisk partners with OpenAI to fasten drug development SPNI acquires TV and digital rights for Indian Football League Qlik partners with ServiceNow to enhance AI-driven enterprise workflows Anthropic hires Trump-linked lobbying firm Ballard Partners OpenAI's $852 billion valuation faces investor scrutiny amid strategy shift Sify data centre arm IPO on track and will be timed with market conditions, says CFO Tata Group asks TCS COO to investigate Nashik sexual harassment case
Popular enterprise AI tools fail to accurately transcribe...
2026-05-11 · via Latest InfoTech News, IT, Information Technology News | The HinduBusinessLine

Popular AI models fail to effectively transcribe Indic languages, mishearing one in three words or dropping English words altogether in mixed speech, as per a study by physical and voice AI data infrastructure company Humyn Labs.

Founded by gaming veteran Manish Agarwal, the startup looks to create a Benchmark of Regional & International Data for Global Evaluation (BRIDGE) to evaluate commercial AI speech-recognition tools. The study looked at tools like ElevenLabs Scribe v2, Deepgram Nova-3, Gemini 2.5 Flash, OpenAI GPT-4o, and Indian providers Sarvam saaras v3 and Gnani vachana v3 on real Indian language data.

The study showed that even the most widely deployed tools have a fundamental problem of mishearing words in Indian language audio. Worse still, in cases of a natural mixing of Hindi or any Indic language with English mid-sentence, most AI tools either drop the English words or convert them into transliterated script, breaking the meaning for anyone reading the transcript.

“The models are grading their own work. ASR providers published their own accuracy scores using benchmarks built on English-first, internet-trained datasets, with little independent validation. Meanwhile, enterprises are making million-dollar deployment decisions on numbers that rarely reflect how their users in Global South actually speak,” said Manish Agarwal, Co-founder, Humyn Labs, adding that theirs is the first independent benchmark for real-world conversational audio across non-English markets.

The scores reveal that Deepgram Nova-3 leads in terms of the semantic gap at 0.906. Amazon Transcribe scores 0.199. OpenAI’s models fall below 0.4. Most enterprises using these tools were unaware of the errors because the standard industry measure, Word Error Rate (WER), was never designed to catch the failures that define real Indian speech.

Comparing global models against Indian providers, the study showed that Sarvam AI’s saaras v3 ranks third overall on WER at 20.2 per cent, ahead of Google Gemini, Microsoft Azure, and AWS Transcribe, a strong result for a model built specifically for Indian languages. However, in terms of mixed speech, Sarvam scores 0.588, placing it in the partial-reliability category where performance varies by language and English density. This means the gap between headline accuracy and code-switch reliability applies to domestic and international providers alike.

Humyn applies a seven-metric stack to test whether the AI models preserve the meaning of what was said, ensure the LLM accurately tracks English words embedded in Indian language speech, how Indic phonology is transcribed as well as Word Information Lost in case of under- or over-transcription.

“The models aren’t the only problem the metrics are. You cannot evaluate non-English speech with a scoring system designed for English phonology and call it rigorous. The performance leaderboard for Hindi is not the leaderboard for Tamil, Bengali and Marathi. A single aggregate benchmark score cannot support cross-regional deployment decisions,” said Ishank Gupta, Co-founder, Humyn Labs.

The study highlights how a model that leads on Spanish may not lead on Vietnamese. Similarly, the model that leads on code-switching does not lead on word accuracy, stressing the need for enterprises to evaluate the language, dialect, and speech pattern that matches their actual users.

Published on May 11, 2026