惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

爱范儿
爱范儿
大猫的无限游戏
大猫的无限游戏
J
Java Code Geeks
MongoDB | Blog
MongoDB | Blog
Martin Fowler
Martin Fowler
GbyAI
GbyAI
Microsoft Azure Blog
Microsoft Azure Blog
Recent Announcements
Recent Announcements
F
Fortinet All Blogs
B
Blog
U
Unit 42
B
Blog RSS Feed
D
DataBreaches.Net
Google DeepMind News
Google DeepMind News
人人都是产品经理
人人都是产品经理
腾讯CDC
量子位
酷 壳 – CoolShell
酷 壳 – CoolShell
V
Visual Studio Blog
博客园 - 聂微东
MyScale Blog
MyScale Blog
奇客Solidot–传递最新科技情报
奇客Solidot–传递最新科技情报
博客园 - 三生石上(FineUI控件)
Engineering at Meta
Engineering at Meta

Business News Today: Latest Business News, Finance News

Markets’ dilemma: Trust the bark or wag of oil prices The sector call illusion Bandu’s Blockbusters For April 12, 2026 Mastering Derivatives: Does Lag Impact Effectiveness Of OI? Who Am I? April 12, 2026 Index Outlook: Rising From Dire Straits US Market Outlook: Gaining Strength Bullion Cues: Gold And Silver Futures Face Barrier F&O Tracker: Tentative Shift In Trend F&O Strategy: Buy L&T Put Maruti Suzuki to launch 4 EVs by 2031 India Inc flags surge in cost of packaging raw material, seeks relief measures India-flagged LPG tanker Jag Vikram crosses Strait of Hormuz after US-Iran ceasefire Muted pricing power, rising costs to curb benefits of demand in cement sector: HDFC Securities Iran's new supreme leader Mojtaba Khamenei has severe and disfiguring wounds, sources say No road tax, registration fees for electric vehicles priced up to ₹30 lakh till March 2030: Delhi’s draft EV policy Central Railway to run four special local trains for Ambedkar Jayanti West Asia tensions push up costs for India; further impact hinges on stability: Report ED initiates fresh raids against former Bengal minister Chatterjee in teacher recruitment scam Election Commission reverses Mittal’s DVAC posting, appoints him DGP, TN Armed Police Israel and Lebanon are expected to hold talks. Here’s what to know US, Iran set for peace talks but doubts emerge over Lebanon, sanctions Cotton Association revises output estimates for 2025-26 up at 324 lakh bales of 170 kg each Orbicular gets USFDA’s tentative nod for generic Semaglutide Injection in partnership with Apotex Malls, high-streets in NCR clock 45% rise in leasing of retail spaces in Jan-Mar: C&W FIIs pull ₹28,375 crore in five sessions; domestic buyers cushion fall as indices post best week in months Nifty and Bank Nifty Prediction for the week 13 Apr’26 to 17 Apr’26 by BL GURU Proposed Trump arch in Washington DC includes winged figure, eagles, lions and gold inscriptions 'Ladakh' replaces 'Jammu and Kashmir' in Aadhaar records for UT residents Misri ends US trip with focus on civil nuclear cooperation and LPG exports
LLM collapse: The danger of training LLMs on AI-generated...
By KV Kurmanath · 2026-06-13 · via Business News Today: Latest Business News, Finance News

What happens when a new generation of large language models (LLMs) are trained on data produced by their predecessors? As internet content becomes increasingly populated with synthetic data, this could lead to a peculiar problem – LLMs can collapse under the weight of AI-generated training data.

Researchers warn it can lead to the ‘collapse of models’ as they are fed data produced by the first-generation LLMs, depending less on the original data. This, in turn, could make later models misperceive reality.

According to a paper published in Nature, using model-generated content in training caused irreversible defects in the resulting models, making the tails of the original content distribution disappear. To sustain AI development in the long term, the authors argue that access to the original data source must be preserved.

“We discover that indiscriminately learning from data produced by other models causes ‘model collapse’,” the authors of the paper, ‘AI models collapse when trained on recursively generated data,’ said.

Raghava Rao Mukkamala, Professor in the Department of Digitalization at Copenhagen Business School, Denmark, explains that current Generative AI models produce new data by learning statistical patterns from massive training datasets, much of which is scraped from the Internet. 

“Unlike humans, who communicate through reasoned intent, experience, and logical argumentation, AI models generate content by applying probabilistic models to patterns they observed in their training data,” he said.

“The Nature paper showed that this recursive training cycle causes models to progressively lose track of the true underlying data distribution in the real world. It showed that rare and diverse patterns are often the first to disappear, making AI outputs increasingly homogeneous, repetitive, and detached from reality,” he told businessline.

Their findings suggest that relying on AI-generated content for future training of AI models may degrade model quality over time. Finally, this study highlights that preserving access to authentic, real-world, and human-generated data is absolutely essential for maintaining the diversity, accuracy, and reliability of future AI systems.

Kashyap Kompella, Chief Executive Officer of RPA2AI Research, said that AI models have been improving because of three main scaling factors: more compute, larger models, and more training data.

“For the last decade, the industry treated web-scale human content as a vast natural resource. That assumption is now breaking. The public availability of high-quality human-generated text is finite. If scaling trends continue, language models could fully use the available stock of public human-generated text by 2032,” he said.

This, however, does not mean “there is no more data.” It means the easiest, cheapest, broadest pool of public human text is no longer enough to keep scaling models in the old way.

“The industry is therefore moving toward licensed data, proprietary enterprise data, human feedback, interaction logs, multimodal data, simulation data, and synthetic data,” he said.

Stating that synthetic data is not automatically bad, it is already useful in code, math, robotics, gaming, autonomous driving, privacy-safe testing, rare-case simulation, and instruction tuning. “The problem begins when synthetic data becomes a substitute for a real-world signal rather than a controlled supplement. Model collapse occurs when AI models are repeatedly trained on outputs from earlier models,” he said.

Who will have an edge?

For the AI Vendors, data quality becomes a strategic moat. Vendors with access to licensed archives, proprietary usage data, enterprise data, multimodal streams, and verified human feedback will have an advantage over vendors relying mainly on public web crawls.

Users will encounter more polished but less original content. The internet will contain more content that looks clear, formatted, and confident but is derivative, repetitive, or weakly sourced. 

Published on June 12, 2026