惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

P
Proofpoint News Feed
Martin Fowler
Martin Fowler
The GitHub Blog
The GitHub Blog
B
Blog RSS Feed
U
Unit 42
阮一峰的网络日志
阮一峰的网络日志
量子位
GbyAI
GbyAI
Cyber Security Advisories - MS-ISAC
Cyber Security Advisories - MS-ISAC
云风的 BLOG
云风的 BLOG
小众软件
小众软件
博客园 - 三生石上(FineUI控件)
L
LangChain Blog
奇客Solidot–传递最新科技情报
奇客Solidot–传递最新科技情报
博客园_首页
IT之家
IT之家
V
Visual Studio Blog
Y
Y Combinator Blog
Blog — PlanetScale
Blog — PlanetScale
宝玉的分享
宝玉的分享
Apple Machine Learning Research
Apple Machine Learning Research
I
InfoQ
D
Docker
V
V2EX

Latest InfoTech News, IT, Information Technology News | The HinduBusinessLine

Google debuts standalone Gemini App for Apple’s MacOS India’s electronics imports cross $116 billion in FY26, exports trail Labour Ministry to look into POSH compliance by IT services firms, says employees union Is TCS harassment case tip of the iceberg? Entry-level smartphones get costlier as memory shortage persists Indians most nervous about AI despite highest skill penetration EdgeCortix secures Axiro, MPower investment to accelerate AI chip innovation Infosys partners Carlos Alcaraz as global brand ambassador Wipro buys select Alpha Net Group contracts for $70.8 Mn AMS expands Pune GCC, strengthens India’s role in global talent operations Memory chip crunch and Iran war lead phone market decline, IDC says UST, Evaaya jointly launch UST Nimbus to help empower GCCs with new capabilities No layoffs, says Zoho: 300 mentioned in social media post were interns Nvidia’s New AI models spark rally in quantum computing stocks OpenAI unveils GPT-5.4-Cyber a week after rival's announcement of AI model IMF urges nations to stay at frontier of mounting AI risks How your CCTV becomes a hacker’s spy Vehant Technologies eyes 20% topline from export in 3 years Cabinet Secretary emphasizes AI development and civil-military cooperation Amazon to acquire Globalstar for $11.57 billion to boost satellite internet Kaar Tech eyes data analytics acquisition as it positions itself as an AI-led enterprise OS enabler India’s quantum mission to complete 2,000 km network by 2027 Andhra Pradesh launches India’s first quantum reference facility in Amaravati Wegovy-maker Novo Nordisk partners with OpenAI to fasten drug development SPNI acquires TV and digital rights for Indian Football League Qlik partners with ServiceNow to enhance AI-driven enterprise workflows Anthropic hires Trump-linked lobbying firm Ballard Partners OpenAI's $852 billion valuation faces investor scrutiny amid strategy shift Sify data centre arm IPO on track and will be timed with market conditions, says CFO Tata Group asks TCS COO to investigate Nashik sexual harassment case
LLM collapse: The danger of training LLMs on AI-generated...
By KV Kurmanath · 2026-06-13 · via Latest InfoTech News, IT, Information Technology News | The HinduBusinessLine

What happens when a new generation of large language models (LLMs) are trained on data produced by their predecessors? As internet content becomes increasingly populated with synthetic data, this could lead to a peculiar problem – LLMs can collapse under the weight of AI-generated training data.

Researchers warn it can lead to the ‘collapse of models’ as they are fed data produced by the first-generation LLMs, depending less on the original data. This, in turn, could make later models misperceive reality.

According to a paper published in Nature, using model-generated content in training caused irreversible defects in the resulting models, making the tails of the original content distribution disappear. To sustain AI development in the long term, the authors argue that access to the original data source must be preserved.

“We discover that indiscriminately learning from data produced by other models causes ‘model collapse’,” the authors of the paper, ‘AI models collapse when trained on recursively generated data,’ said.

Raghava Rao Mukkamala, Professor in the Department of Digitalization at Copenhagen Business School, Denmark, explains that current Generative AI models produce new data by learning statistical patterns from massive training datasets, much of which is scraped from the Internet. 

“Unlike humans, who communicate through reasoned intent, experience, and logical argumentation, AI models generate content by applying probabilistic models to patterns they observed in their training data,” he said.

“The Nature paper showed that this recursive training cycle causes models to progressively lose track of the true underlying data distribution in the real world. It showed that rare and diverse patterns are often the first to disappear, making AI outputs increasingly homogeneous, repetitive, and detached from reality,” he told businessline.

Their findings suggest that relying on AI-generated content for future training of AI models may degrade model quality over time. Finally, this study highlights that preserving access to authentic, real-world, and human-generated data is absolutely essential for maintaining the diversity, accuracy, and reliability of future AI systems.

Kashyap Kompella, Chief Executive Officer of RPA2AI Research, said that AI models have been improving because of three main scaling factors: more compute, larger models, and more training data.

“For the last decade, the industry treated web-scale human content as a vast natural resource. That assumption is now breaking. The public availability of high-quality human-generated text is finite. If scaling trends continue, language models could fully use the available stock of public human-generated text by 2032,” he said.

This, however, does not mean “there is no more data.” It means the easiest, cheapest, broadest pool of public human text is no longer enough to keep scaling models in the old way.

“The industry is therefore moving toward licensed data, proprietary enterprise data, human feedback, interaction logs, multimodal data, simulation data, and synthetic data,” he said.

Stating that synthetic data is not automatically bad, it is already useful in code, math, robotics, gaming, autonomous driving, privacy-safe testing, rare-case simulation, and instruction tuning. “The problem begins when synthetic data becomes a substitute for a real-world signal rather than a controlled supplement. Model collapse occurs when AI models are repeatedly trained on outputs from earlier models,” he said.

Who will have an edge?

For the AI Vendors, data quality becomes a strategic moat. Vendors with access to licensed archives, proprietary usage data, enterprise data, multimodal streams, and verified human feedback will have an advantage over vendors relying mainly on public web crawls.

Users will encounter more polished but less original content. The internet will contain more content that looks clear, formatted, and confident but is derivative, repetitive, or weakly sourced. 

Published on June 12, 2026