惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

雷峰网
雷峰网
MongoDB | Blog
MongoDB | Blog
D
Docker
Martin Fowler
Martin Fowler
人人都是产品经理
人人都是产品经理
GbyAI
GbyAI
Jina AI
Jina AI
酷 壳 – CoolShell
酷 壳 – CoolShell
M
MIT News - Artificial intelligence
腾讯CDC
阮一峰的网络日志
阮一峰的网络日志
H
Hackread – Cybersecurity News, Data Breaches, AI and More
N
Netflix TechBlog - Medium
B
Blog RSS Feed
云风的 BLOG
云风的 BLOG
Blog — PlanetScale
Blog — PlanetScale
Vercel News
Vercel News
The Cloudflare Blog
freeCodeCamp Programming Tutorials: Python, JavaScript, Git & More
有赞技术团队
有赞技术团队
G
Google Developers Blog
Stack Overflow Blog
Stack Overflow Blog
I
InfoQ
U
Unit 42

Latest InfoTech News, IT, Information Technology News | The HinduBusinessLine

Google debuts standalone Gemini App for Apple’s MacOS India’s electronics imports cross $116 billion in FY26, exports trail Labour Ministry to look into POSH compliance by IT services firms, says employees union Is TCS harassment case tip of the iceberg? Entry-level smartphones get costlier as memory shortage persists Indians most nervous about AI despite highest skill penetration EdgeCortix secures Axiro, MPower investment to accelerate AI chip innovation Infosys partners Carlos Alcaraz as global brand ambassador Wipro buys select Alpha Net Group contracts for $70.8 Mn AMS expands Pune GCC, strengthens India’s role in global talent operations Memory chip crunch and Iran war lead phone market decline, IDC says UST, Evaaya jointly launch UST Nimbus to help empower GCCs with new capabilities No layoffs, says Zoho: 300 mentioned in social media post were interns Nvidia’s New AI models spark rally in quantum computing stocks OpenAI unveils GPT-5.4-Cyber a week after rival's announcement of AI model IMF urges nations to stay at frontier of mounting AI risks How your CCTV becomes a hacker’s spy Vehant Technologies eyes 20% topline from export in 3 years Cabinet Secretary emphasizes AI development and civil-military cooperation Amazon to acquire Globalstar for $11.57 billion to boost satellite internet Kaar Tech eyes data analytics acquisition as it positions itself as an AI-led enterprise OS enabler India’s quantum mission to complete 2,000 km network by 2027 Andhra Pradesh launches India’s first quantum reference facility in Amaravati Wegovy-maker Novo Nordisk partners with OpenAI to fasten drug development SPNI acquires TV and digital rights for Indian Football League Qlik partners with ServiceNow to enhance AI-driven enterprise workflows Anthropic hires Trump-linked lobbying firm Ballard Partners OpenAI's $852 billion valuation faces investor scrutiny amid strategy shift Sify data centre arm IPO on track and will be timed with market conditions, says CFO Tata Group asks TCS COO to investigate Nashik sexual harassment case
Popular enterprise AI tools fail to accurately transcribe...
2026-05-11 · via Latest InfoTech News, IT, Information Technology News | The HinduBusinessLine

Popular AI models fail to effectively transcribe Indic languages, mishearing one in three words or dropping English words altogether in mixed speech, as per a study by physical and voice AI data infrastructure company Humyn Labs.

Founded by gaming veteran Manish Agarwal, the startup looks to create a Benchmark of Regional & International Data for Global Evaluation (BRIDGE) to evaluate commercial AI speech-recognition tools. The study looked at tools like ElevenLabs Scribe v2, Deepgram Nova-3, Gemini 2.5 Flash, OpenAI GPT-4o, and Indian providers Sarvam saaras v3 and Gnani vachana v3 on real Indian language data.

The study showed that even the most widely deployed tools have a fundamental problem of mishearing words in Indian language audio. Worse still, in cases of a natural mixing of Hindi or any Indic language with English mid-sentence, most AI tools either drop the English words or convert them into transliterated script, breaking the meaning for anyone reading the transcript.

“The models are grading their own work. ASR providers published their own accuracy scores using benchmarks built on English-first, internet-trained datasets, with little independent validation. Meanwhile, enterprises are making million-dollar deployment decisions on numbers that rarely reflect how their users in Global South actually speak,” said Manish Agarwal, Co-founder, Humyn Labs, adding that theirs is the first independent benchmark for real-world conversational audio across non-English markets.

The scores reveal that Deepgram Nova-3 leads in terms of the semantic gap at 0.906. Amazon Transcribe scores 0.199. OpenAI’s models fall below 0.4. Most enterprises using these tools were unaware of the errors because the standard industry measure, Word Error Rate (WER), was never designed to catch the failures that define real Indian speech.

Comparing global models against Indian providers, the study showed that Sarvam AI’s saaras v3 ranks third overall on WER at 20.2 per cent, ahead of Google Gemini, Microsoft Azure, and AWS Transcribe, a strong result for a model built specifically for Indian languages. However, in terms of mixed speech, Sarvam scores 0.588, placing it in the partial-reliability category where performance varies by language and English density. This means the gap between headline accuracy and code-switch reliability applies to domestic and international providers alike.

Humyn applies a seven-metric stack to test whether the AI models preserve the meaning of what was said, ensure the LLM accurately tracks English words embedded in Indian language speech, how Indic phonology is transcribed as well as Word Information Lost in case of under- or over-transcription.

“The models aren’t the only problem the metrics are. You cannot evaluate non-English speech with a scoring system designed for English phonology and call it rigorous. The performance leaderboard for Hindi is not the leaderboard for Tamil, Bengali and Marathi. A single aggregate benchmark score cannot support cross-regional deployment decisions,” said Ishank Gupta, Co-founder, Humyn Labs.

The study highlights how a model that leads on Spanish may not lead on Vietnamese. Similarly, the model that leads on code-switching does not lead on word accuracy, stressing the need for enterprises to evaluate the language, dialect, and speech pattern that matches their actual users.

Published on May 11, 2026