惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

Microsoft Security Blog
Microsoft Security Blog
Apple Machine Learning Research
Apple Machine Learning Research
美团技术团队
WordPress大学
WordPress大学
酷 壳 – CoolShell
酷 壳 – CoolShell
G
Google Developers Blog
阮一峰的网络日志
阮一峰的网络日志
The Cloudflare Blog
J
Java Code Geeks
Martin Fowler
Martin Fowler
M
MIT News - Artificial intelligence
IT之家
IT之家
博客园 - 三生石上(FineUI控件)
月光博客
月光博客
Google DeepMind News
Google DeepMind News
小众软件
小众软件
V
V2EX
Hugging Face - Blog
Hugging Face - Blog
爱范儿
爱范儿
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻
Jina AI
Jina AI
freeCodeCamp Programming Tutorials: Python, JavaScript, Git & More
腾讯CDC
B
Blog

informationweek

2026 tech company layoffs How Sedgwick scaled AI in legacy claims workflows InformationWeek Podcast: CTOs on using AI in regulated spaces How top CIOs are measuring the real ROI of IT automation What AI must learn from Roosevelt, conservation and 1929 Experian's chief innovation officer gleans AI gains with startup collab ETS CIO on competing with AI startups 'running with scissors' Before the next VMware: How CIOs prepare for vendor shocks The strategic alignment powering cyber-resilient organizations The AI infrastructure bottleneck is becoming a CIO problem InformationWeek Podcast: CTOs on reining in rogue AI agents Workplace equity in the age of AI Why and how to implement an AI asset rationalization strategy Why companies are shifting toward private AI models AI agents in automation: When to build, when to buy Navan CTO AI on trial: The Workday case that CIOs can The AI infrastructure boom is coming for enterprise budgets How CIOs can manage LLM costs: A practical guide What CIOs miss when buying vertical SaaS software InformationWeek Podcast: How CTOs balance AI and their teams Whirlpool, Duke Energy, Cleveland Clinic CIOs on scaling AI Where CIOs get stuck rebuilding the enterprise: What 'Rewired' reveals As AI makes projects harder to track, will CIOs need new controls? Why disaster recovery plans fail in geopolitical crises A silent erosion of enterprise AI by data poisoning Priceline CTO prioritizes engineers able to 'hold a room and a roadmap' InformationWeek Podcast: When CTOs need to restart IT projects Wayfair CTO maps agentic path across digital and brick-and-mortar commerce The AI contract gaps the Google-Pentagon deal just made visible
Why AI teams treat training data like capital
Daniel Mande · 2026-04-21 · via informationweek

Early artificial intelligence development operated on an assumption: Data was abundant, and -- if not exactly free -- it was at least treated as a low-friction input. Compute was scarce. Talent was scarce. GPUs had line items. Data, by contrast, was scraped or acquired and absorbed into models, often with limited documentation of provenance, structured metadata or niche data to support long-term reuse.

That era is ending.

Model builders are now evaluating data the way teams evaluate infrastructure investments or capital expenditures: by pricing legal risk and quality, and accounting for future optionality. 

Historically, data costs were real but indirect. A team might pay for a data set or scrape public web content. The expense appeared as a one-time acquisition cost or as a line item buried in operating budgets. Once ingested into a model, the data largely disappeared from view, even as it continued to shape downstream products, performance and risk.

Related:The AI revolution: We've seen this movie before

Litigation risk was often treated as theoretical. Regulatory requirements around training data were ambiguous or nonexistent. As long as models performed well and revenue grew, few organizations revisited the provenance of the data embedded inside their systems.

A shift began when litigation moved from speculative to concrete. Cases have signaled that courts are willing to scrutinize how AI companies acquire and use proprietary content. Regardless of how individual cases resolve, the mere fact that they exist changes the calculus.

Regulation is operationalizing what was once theoretical, and regulators are pushing for greater transparency into training data sources and governance. 

This creates exposure if a company cannot clearly document what went into its model, including rights status, licensing terms and data provenance. If those inputs are later challenged, the cost isn’t confined to the budget. It can manifest as delayed deployments, constrained market access, forced model retraining or reputational damage.

Economic consequences are already here

The financial impact of poor data decisions is real. Incomplete, too generalized or biased data sets can degrade model performance in ways that are expensive and difficult to reverse. As AI systems become more embedded in revenue-generating workflows, the cost of flawed or contested data compounds. The impact shows up in not just research metrics, but also balance sheets.

Data decisions now have enterprise-level consequences, and those consequences can no longer be deferred.

Related:How Ethical Scorecards Help Build Trust in AI Systems

From input to asset

When an input creates long-lived exposure and long-lived value, it begins to look like capital.

Training data increasingly fits that description. A continuously refreshed, high-quality, labeled and domain-specific corpus can be reused across models, geographies and product lines. It can accelerate compliance. It can shorten procurement cycles with enterprise customers who demand provenance clarity. It can serve as a defensible moat.

Conversely, poorly governed data accumulates hidden liabilities. If a data set’s legal status is uncertain, its downstream uses may be constrained. If documentation is incomplete, audit costs rise. If rights are ambiguous, partnerships stall.

AI teams are starting to recognize this dynamic. They are modeling not just the immediate performance gains from adding a data set, but also the lifecycle implications: Can this data be reused across multiple model generations? Does it increase or decrease regulatory friction? What is the expected cost of litigation or forced retraining? 

These are capital allocation questions.

The counterargument: Fair use will hold

Not everyone accepts this framing. Some AI teams continue to operate under the assumption that broad fair-use interpretations will remain viable and that large-scale web scraping will ultimately be vindicated in court. 

Related:Is MCP the key to unlocking autonomous enterprise AI?

There is a rational logic here. Courts may indeed affirm expansive interpretations of fair use in certain contexts. Regulatory enforcement may evolve slowly.

But this argument underestimates a critical factor: uncertainty itself carries cost.

Uncertainty narrows optionality. If a model’s training data is legally ambiguous, a company may avoid expanding into regulated markets, or it may hesitate to retrain or fine-tune in ways that could trigger fresh scrutiny.

A capital discipline for data

Treating data like capital does not mean slowing innovation. It means building on a stronger foundation.

Capital investments are evaluated for durability, return and risk exposure. Training data increasingly deserves the same scrutiny. Rights-cleared, multimodal data sets with strong provenance reduce legal uncertainty, improve model performance, accelerate enterprise adoption and preserve long-term optionality.

About the Author

Daniel Mandell

Shutterstock

Daniel Mandell is senior vice president of data licensing and AI services at Shutterstock, with more than 20 years of experience building and scaling digital media, data and AI-driven businesses.

Daniel leads a team focused on unlocking the value of data through strategic licensing, data services and AI-powered solutions for AI companies and model builders.