惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

J
Java Code Geeks
Stack Overflow Blog
Stack Overflow Blog
B
Blog RSS Feed
C
Check Point Blog
D
Docker
Y
Y Combinator Blog
Recent Announcements
Recent Announcements
Google DeepMind News
Google DeepMind News
MongoDB | Blog
MongoDB | Blog
博客园_首页
Apple Machine Learning Research
Apple Machine Learning Research
量子位
有赞技术团队
有赞技术团队
IT之家
IT之家
大猫的无限游戏
大猫的无限游戏
D
DataBreaches.Net
M
MIT News - Artificial intelligence
B
Blog
阮一峰的网络日志
阮一峰的网络日志
freeCodeCamp Programming Tutorials: Python, JavaScript, Git & More
腾讯CDC
让小产品的独立变现更简单 - ezindie.com
让小产品的独立变现更简单 - ezindie.com
V
V2EX
月光博客
月光博客

informationweek

2026 tech company layoffs How Sedgwick scaled AI in legacy claims workflows InformationWeek Podcast: CTOs on using AI in regulated spaces How top CIOs are measuring the real ROI of IT automation What AI must learn from Roosevelt, conservation and 1929 Experian's chief innovation officer gleans AI gains with startup collab ETS CIO on competing with AI startups 'running with scissors' Before the next VMware: How CIOs prepare for vendor shocks The strategic alignment powering cyber-resilient organizations The AI infrastructure bottleneck is becoming a CIO problem InformationWeek Podcast: CTOs on reining in rogue AI agents Workplace equity in the age of AI Why and how to implement an AI asset rationalization strategy Why companies are shifting toward private AI models AI agents in automation: When to build, when to buy Navan CTO AI on trial: The Workday case that CIOs can The AI infrastructure boom is coming for enterprise budgets How CIOs can manage LLM costs: A practical guide What CIOs miss when buying vertical SaaS software InformationWeek Podcast: How CTOs balance AI and their teams Whirlpool, Duke Energy, Cleveland Clinic CIOs on scaling AI Where CIOs get stuck rebuilding the enterprise: What 'Rewired' reveals As AI makes projects harder to track, will CIOs need new controls? Why disaster recovery plans fail in geopolitical crises Priceline CTO prioritizes engineers able to 'hold a room and a roadmap' InformationWeek Podcast: When CTOs need to restart IT projects Wayfair CTO maps agentic path across digital and brick-and-mortar commerce The AI contract gaps the Google-Pentagon deal just made visible Non-human identity sprawl is agentic AI's real risk
A silent erosion of enterprise AI by data poisoning
Niranjan Krishnan · 2026-05-05 · via informationweek

Image of a human brain, data processing and network of connections

Wavebreakmedia Ltd IFE-240314_16/Alamy

When big data went mainstream a decade ago, data lakes were filled with insights, patterns and predictions driven by machine learning. Quality improved over time as automated data collection enriched training data sets, and feedback loops enabled rapid retraining. 

The result was a virtuous cycle of better data, better models and better decisions.

A similar phenomenon is emerging in generative AI, but in reverse.

As enterprises deploy AI across business functions, data environments are being inundated with synthetic content, such as summaries, emails, reports, code and images. While synthetic data can be valuable when real-world data is unavailable, ambient AI-generated content introduces a more systemic risk: inadvertent data poisoning.

Unlike traditional data poisoning in cybersecurity, this isn't malicious. It's self-inflicted, but no less damaging.

The death spiral of recursive training

AI models learn from abstractions of the real world. When training data drifts away from first-hand reality, models begin to learn from their own approximations rather than facts. Over time, they lose the ability to distinguish truth from statistical likelihood.

Related:Intuit's chief AI officer on the SaaSpocalypse and disciplined AI

A feedback loop accelerates this process. With each iteration, models smooth out edge cases and converge toward safer, more generic outputs. While this may work for common scenarios, it can create risk in rare but critical situations.

Consider how engineers design dams. A dam built for average rainfall will perform most of the time, but it can fail catastrophically during a 100-year flood. Similarly, models trained on AI-generated data may perform adequately in routine cases but break down under stress, when nuance and precision matter most.

Hallucinated content compounds the problem, introducing errors that are then reinforced through retraining.

The impact is gradual but significant: Outputs become less precise and less diverse, and they are less grounded in reality. This is the early stage of what researchers call "model collapse."

The math of model collapse

A 2024 paper in Nature by Shumailov et al. formalized "model collapse," showing that training on AI-generated data leads to irreversible performance degradation. As models retrain on their own outputs, they effectively trim the "tails" of the data distribution, the very areas where rare but high-value insights exist.

The result is regression to the mean: a loss of nuance, diversity and real-world fidelity.

A simple analogy is photocopying a document repeatedly. Each copy loses detail until only the broad outlines remain. In the same way, AI systems trained on degraded data lose the fidelity required to support complex business decisions.

Related:Time for an AI exit strategy: How CIOs are cutting AI waste

The compliance trap

This erosion also amplifies algorithmic bias. AI models already reflect patterns in their training data. When trained on AI-generated content, those biases are reinforced and magnified. The result is not just degraded performance but also increased regulatory and compliance risk.

Once a model collapses, no amount of fine-tuning can restore it. The only solution is disciplined data governance.

Organizations should take several steps:

  • Manage data as products, with lifecycle controls and quality standards.

  • Exclude AI-generated content by default from training pipelines.

  • Establish data provenance, using techniques like watermarking to track data's origin.

  • Tag data at ingestion as AI-generated, AI-edited or original.

  • Invest in "golden data sets" to anchor models in real-world truth.

These practices ensure that training data remains grounded, traceable and fit for purpose.

The new competitive edge

A longstanding principle in data science still holds: Clean data beats clever algorithms.

In today's AI landscape, this is no longer a best practice; it is a competitive necessity. As models and tools commoditize, they cease to differentiate. High-quality, well-governed data becomes the only durable advantage.

Related:CIOs need control before AI gains accountability

Organizations that allow AI-generated content to flow unchecked into their data ecosystems are not just introducing noise; they are also eroding the very foundation of their AI capabilities.

The winners will not be those with the most data, but those with the cleanest, most human-centric data.

About the Author

Niranjan Krishnan

FPT Americas

Niranjan Krishnan is head of AI solutions at FPT Americas with two decades of experience in delivering on the promise of data.

Niranjan has led large cross-functional teams and deployed dozens of AI/machine learning solutions across industries. He is passionate about responsible AI solutions that create measurable value for businesses and customers.