惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

U
Unit 42
Vercel News
Vercel News
博客园 - 叶小钗
大猫的无限游戏
大猫的无限游戏
MyScale Blog
MyScale Blog
P
Proofpoint News Feed
量子位
Engineering at Meta
Engineering at Meta
B
Blog RSS Feed
博客园 - 【当耐特】
Recent Announcements
Recent Announcements
Google DeepMind News
Google DeepMind News
D
DataBreaches.Net
Stack Overflow Blog
Stack Overflow Blog
博客园 - 聂微东
小众软件
小众软件
Hugging Face - Blog
Hugging Face - Blog
人人都是产品经理
人人都是产品经理
IT之家
IT之家
T
The Blog of Author Tim Ferriss
Last Week in AI
Last Week in AI
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻
Jina AI
Jina AI
博客园 - 三生石上(FineUI控件)

Opinion, Editorial, Views, Columnists, Columns | The HinduBusinessLine

Rupee can’t be defended from just one side Railways’ performance Why not have a women-only party? Labour pangs Pak’s peculiar comeback on the global stage Letters to Editor India has jobs, but it needs better ones Cross-border insolvency laws and trade A major health challenge Editorial. Snooping around Letters to the Editor dated April 20, 2026 All you want to know about the women’s reservation and delimitation bills fiasco Editorial. Process deficit Letters to the Editor dated April 19, 2026 WPI effect on new GDP series The tragic reality of police brutality India’s AI value paradox Prepare the ground India-Korea economic ties poised to strengthen Nari Shakti Bill — a missed opportunity Natural farming should become mainstream policy Insights from new GDP data Strategies to enhance fertilizer security Pathway to maritime insurance sovereignty Why the GoP’s jittery Clear the smoke Aiding piped gas push Stocks are the least over-priced asset in India Is TCS harassment case tip of the iceberg? SIP with caution
Indian languages, the foundation of India’s AI
TV Ramachandran & KV Seshasayee · 2026-06-17 · via Opinion, Editorial, Views, Columnists, Columns | The HinduBusinessLine
AI: Powered on Indian languages

AI: Powered on Indian languages | Photo Credit: Rawf8

India’s ambitions in artificial intelligence are growing rapidly. Governments are investing in AI infrastructure, startups are attracting capital, and research institutions are building increasingly capable language models. Yet one fundamental challenge remains largely overlooked: India’s sovereign AI goal cannot be met without a strong knowledge infrastructure for its own languages.

Today, AI systems perform best when trained on large volumes of high-quality digital content. For English, such content exists in abundance. For most Indian languages, it does not. This is emerging as the single biggest bottleneck in the development of truly inclusive and effective Indian-language AI.

While Hindi enjoys a relatively rich digital footprint, languages such as Tamil, Telugu, Bengali and Marathi have more limited resources, and many others remain severely underrepresented. As a result, AI systems often struggle with accuracy, reasoning, summarisation and translation in these languages. The problem is not primarily one of computing power or model architecture. It is the lack of clean, diverse and digitised text that reflects India’s linguistic and cultural richness.

Beyond tech

The implications extend far beyond technology. Large-scale digitisation is essential for modern governance, education, legal systems and cultural preservation. Government records, court judgments, land documents, textbooks, research papers and historical archives all need to become machine-readable if AI is to deliver meaningful public value.

At the heart of this challenge lies Optical Character Recognition (OCR) — the technology that converts scanned documents into searchable and usable text. OCR is often taken for granted in English, but for many Indian scripts it remains a significant hurdle.

Even printed documents present difficulties. Government records are frequently available only as low-quality scanned PDFs. Newspapers and books often use non-standard fonts. Complex page layouts containing multiple columns, tables and scripts reduce accuracy further. For languages such as Tamil, Malayalam and Urdu, OCR performance remains uneven.

The challenge becomes even greater when dealing with handwritten material. Millions of government records, historical archives and institutional documents remain handwritten. Regional variations in handwriting, the absence of large labelled datasets and older writing styles make automated recognition extremely difficult.

India’s vast manuscript heritage presents another frontier. Palm-leaf manuscripts, copper-plate inscriptions and ancient texts contain centuries of knowledge in fields ranging from mathematics and astronomy to medicine and philosophy. Unlocking these resources requires not only OCR but also image restoration, script identification and linguistic expertise. This is as much a national knowledge mission as a technology project.

Indian initiatives

Fortunately, important foundations already exist.

AI4Bharat at IIT Madras has emerged as one of India’s most significant open-source initiatives, contributing multilingual datasets, evaluation benchmarks and language models for Indian languages. IIIT Hyderabad has undertaken important work in OCR and document analysis. Meanwhile, companies such as Sarvam AI, BharatGPT, Microsoft, Google and Meta are investing in deployment and innovation.

Government initiatives have also made notable contributions. Bhashini has advanced speech and translation technologies. The National Manuscripts Mission has surveyed millions of manuscripts. The National Digital Library has assembled a large collection of digital resources. Several states, including Tamil Nadu, Kerala and Karnataka, have launched valuable digitisation programmes.

Yet these efforts remain fragmented. India still lacks common standards, interoperable datasets, AI-ready pipelines and a coordinated national strategy. The result is duplication of effort and slower progress than the country requires.

What India needs now is a National Knowledge Infrastructure for Indic AI.

First, a National Text Recognition Mission should be launched to accelerate development of OCR systems, handwriting recognition technologies, manuscript digitisation capabilities and next-generation vision-language models tailored to Indian scripts.

Second, a National Corpus Authority should establish standards for metadata, data quality, storage and interoperability while coordinating contributions from governments, universities, libraries and cultural institutions.

Third, India requires a modern licensing framework that balances public access, intellectual property protection and fair compensation for publishers and content creators.

Finally, stronger collaboration between government, academia, industry and civil society is essential. India’s linguistic diversity is unmatched globally. No single institution can solve this challenge alone.

As Nandan Nilekani recently argued, India has already shown through Digital Public Infrastructure such as UPI how open, interoperable public platforms can create transformative national outcomes. AI can follow a similar path.

The real race in AI is not merely about building bigger models or acquiring more GPUs. It is about creating the knowledge foundations on which those models can learn. If India succeeds in digitising, organising and democratizing access to its linguistic wealth, it can build AI systems that serve not only English-speaking elites but also the hundreds of millions who communicate in Indian languages every day.

The future of Indian AI will clearly epend on how effectively we unlock India’s knowledge treasure-house and make it accessible to machines — and to people.

Seshasayee is Principal Adviser and Ramachandran is President of BIF. Views expressed are personal

Published on June 18, 2026