惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

Vercel News
Vercel News
奇客Solidot–传递最新科技情报
奇客Solidot–传递最新科技情报
Apple Machine Learning Research
Apple Machine Learning Research
T
Tailwind CSS Blog
让小产品的独立变现更简单 - ezindie.com
让小产品的独立变现更简单 - ezindie.com
人人都是产品经理
人人都是产品经理
V
V2EX
量子位
Last Week in AI
Last Week in AI
Jina AI
Jina AI
博客园 - 【当耐特】
爱范儿
爱范儿
宝玉的分享
宝玉的分享
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻
Hugging Face - Blog
Hugging Face - Blog
博客园 - 三生石上(FineUI控件)
有赞技术团队
有赞技术团队
小众软件
小众软件
IT之家
IT之家
博客园_首页
博客园 - 聂微东
S
SegmentFault 最新的问题
阮一峰的网络日志
阮一峰的网络日志
博客园 - 叶小钗

Business News Today: Latest Business News, Finance News

Markets’ dilemma: Trust the bark or wag of oil prices The sector call illusion Bandu’s Blockbusters For April 12, 2026 Mastering Derivatives: Does Lag Impact Effectiveness Of OI? Who Am I? April 12, 2026 Index Outlook: Rising From Dire Straits US Market Outlook: Gaining Strength Bullion Cues: Gold And Silver Futures Face Barrier F&O Tracker: Tentative Shift In Trend F&O Strategy: Buy L&T Put Maruti Suzuki to launch 4 EVs by 2031 India Inc flags surge in cost of packaging raw material, seeks relief measures India-flagged LPG tanker Jag Vikram crosses Strait of Hormuz after US-Iran ceasefire Muted pricing power, rising costs to curb benefits of demand in cement sector: HDFC Securities Iran's new supreme leader Mojtaba Khamenei has severe and disfiguring wounds, sources say No road tax, registration fees for electric vehicles priced up to ₹30 lakh till March 2030: Delhi’s draft EV policy Central Railway to run four special local trains for Ambedkar Jayanti West Asia tensions push up costs for India; further impact hinges on stability: Report ED initiates fresh raids against former Bengal minister Chatterjee in teacher recruitment scam Election Commission reverses Mittal’s DVAC posting, appoints him DGP, TN Armed Police Israel and Lebanon are expected to hold talks. Here’s what to know US, Iran set for peace talks but doubts emerge over Lebanon, sanctions Cotton Association revises output estimates for 2025-26 up at 324 lakh bales of 170 kg each Orbicular gets USFDA’s tentative nod for generic Semaglutide Injection in partnership with Apotex Malls, high-streets in NCR clock 45% rise in leasing of retail spaces in Jan-Mar: C&W FIIs pull ₹28,375 crore in five sessions; domestic buyers cushion fall as indices post best week in months Nifty and Bank Nifty Prediction for the week 13 Apr’26 to 17 Apr’26 by BL GURU Proposed Trump arch in Washington DC includes winged figure, eagles, lions and gold inscriptions 'Ladakh' replaces 'Jammu and Kashmir' in Aadhaar records for UT residents Misri ends US trip with focus on civil nuclear cooperation and LPG exports
Indian languages, the foundation of India’s AI
TV Ramachandran & KV Seshasayee · 2026-06-17 · via Business News Today: Latest Business News, Finance News
AI: Powered on Indian languages

AI: Powered on Indian languages | Photo Credit: Rawf8

India’s ambitions in artificial intelligence are growing rapidly. Governments are investing in AI infrastructure, startups are attracting capital, and research institutions are building increasingly capable language models. Yet one fundamental challenge remains largely overlooked: India’s sovereign AI goal cannot be met without a strong knowledge infrastructure for its own languages.

Today, AI systems perform best when trained on large volumes of high-quality digital content. For English, such content exists in abundance. For most Indian languages, it does not. This is emerging as the single biggest bottleneck in the development of truly inclusive and effective Indian-language AI.

While Hindi enjoys a relatively rich digital footprint, languages such as Tamil, Telugu, Bengali and Marathi have more limited resources, and many others remain severely underrepresented. As a result, AI systems often struggle with accuracy, reasoning, summarisation and translation in these languages. The problem is not primarily one of computing power or model architecture. It is the lack of clean, diverse and digitised text that reflects India’s linguistic and cultural richness.

Beyond tech

The implications extend far beyond technology. Large-scale digitisation is essential for modern governance, education, legal systems and cultural preservation. Government records, court judgments, land documents, textbooks, research papers and historical archives all need to become machine-readable if AI is to deliver meaningful public value.

At the heart of this challenge lies Optical Character Recognition (OCR) — the technology that converts scanned documents into searchable and usable text. OCR is often taken for granted in English, but for many Indian scripts it remains a significant hurdle.

Even printed documents present difficulties. Government records are frequently available only as low-quality scanned PDFs. Newspapers and books often use non-standard fonts. Complex page layouts containing multiple columns, tables and scripts reduce accuracy further. For languages such as Tamil, Malayalam and Urdu, OCR performance remains uneven.

The challenge becomes even greater when dealing with handwritten material. Millions of government records, historical archives and institutional documents remain handwritten. Regional variations in handwriting, the absence of large labelled datasets and older writing styles make automated recognition extremely difficult.

India’s vast manuscript heritage presents another frontier. Palm-leaf manuscripts, copper-plate inscriptions and ancient texts contain centuries of knowledge in fields ranging from mathematics and astronomy to medicine and philosophy. Unlocking these resources requires not only OCR but also image restoration, script identification and linguistic expertise. This is as much a national knowledge mission as a technology project.

Indian initiatives

Fortunately, important foundations already exist.

AI4Bharat at IIT Madras has emerged as one of India’s most significant open-source initiatives, contributing multilingual datasets, evaluation benchmarks and language models for Indian languages. IIIT Hyderabad has undertaken important work in OCR and document analysis. Meanwhile, companies such as Sarvam AI, BharatGPT, Microsoft, Google and Meta are investing in deployment and innovation.

Government initiatives have also made notable contributions. Bhashini has advanced speech and translation technologies. The National Manuscripts Mission has surveyed millions of manuscripts. The National Digital Library has assembled a large collection of digital resources. Several states, including Tamil Nadu, Kerala and Karnataka, have launched valuable digitisation programmes.

Yet these efforts remain fragmented. India still lacks common standards, interoperable datasets, AI-ready pipelines and a coordinated national strategy. The result is duplication of effort and slower progress than the country requires.

What India needs now is a National Knowledge Infrastructure for Indic AI.

First, a National Text Recognition Mission should be launched to accelerate development of OCR systems, handwriting recognition technologies, manuscript digitisation capabilities and next-generation vision-language models tailored to Indian scripts.

Second, a National Corpus Authority should establish standards for metadata, data quality, storage and interoperability while coordinating contributions from governments, universities, libraries and cultural institutions.

Third, India requires a modern licensing framework that balances public access, intellectual property protection and fair compensation for publishers and content creators.

Finally, stronger collaboration between government, academia, industry and civil society is essential. India’s linguistic diversity is unmatched globally. No single institution can solve this challenge alone.

As Nandan Nilekani recently argued, India has already shown through Digital Public Infrastructure such as UPI how open, interoperable public platforms can create transformative national outcomes. AI can follow a similar path.

The real race in AI is not merely about building bigger models or acquiring more GPUs. It is about creating the knowledge foundations on which those models can learn. If India succeeds in digitising, organising and democratizing access to its linguistic wealth, it can build AI systems that serve not only English-speaking elites but also the hundreds of millions who communicate in Indian languages every day.

The future of Indian AI will clearly epend on how effectively we unlock India’s knowledge treasure-house and make it accessible to machines — and to people.

Seshasayee is Principal Adviser and Ramachandran is President of BIF. Views expressed are personal

Published on June 18, 2026