惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

D
DataBreaches.Net
IT之家
IT之家
博客园_首页
博客园 - 【当耐特】
V
V2EX
Apple Machine Learning Research
Apple Machine Learning Research
G
Google Developers Blog
让小产品的独立变现更简单 - ezindie.com
让小产品的独立变现更简单 - ezindie.com
Recent Announcements
Recent Announcements
F
Fortinet All Blogs
GbyAI
GbyAI
腾讯CDC
H
Hackread – Cybersecurity News, Data Breaches, AI and More
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻
I
InfoQ
H
Help Net Security
T
Tailwind CSS Blog
B
Blog RSS Feed
Martin Fowler
Martin Fowler
人人都是产品经理
人人都是产品经理
The Cloudflare Blog
博客园 - 叶小钗
雷峰网
雷峰网
量子位

Forbes - Innovation

Why Do Humans Have Fingerprints? Hint: It’s Not What You Think Booking.com Confirms Data Breach, Reservation PIN Codes Changed Why Major News Sites Are Blocking The Internet Archive’s Wayback Machine iPhone Fold Release Date: New Report Details Frustrating Apple News Comet Tracker: How To See Pan-STARRS And Three Planets On Wednesday NYT Mini Crossword Today: Tuesday, April 14 Hints And Answers Today’s NYT Strands Hints, Spangram, Answers: Tuesday, April 14 (It’s A Little Unclear) Today’s Wordle #1760 Hints And Answer For Tuesday, April 14 Most Of The Microplastics In Urban Air Come From Tires Today’s Wordle #1759 Hints And Answer For Monday, April 13 NYT Mini Crossword Today: Monday, April 13 Hints And Answers NYT Pips Today: Hints, Answers And Walkthrough For Monday, April 13 The YC Chief Who Codes 10,000 Lines A Day Has A Simple Secret Samsung Expands One UI 8.5 Beta To More Galaxy Owners Why You Should Stop Using Your iPhone If It’s On This List Chamath Says Firms That Treat AI As A Strategy Hand Rivals Their Edge 3 Unexpected Habits Of Secure Couples, By A Psychologist The First Lamp That Folds Your Clothes Samsung’s Disappointing Price Update For Galaxy Phone Buyers 3 Subtle Signs Someone Is Falling In Love With You, By A Psychologist Do Mantis Shrimp See More Colors Than Humans? A Biologist Explains NYT Connections Answers Explained For Monday, April 13 (#1,037) NYT Connections Hints Today: Monday, April 13 Clues And Answers (#1,037) LEGO Luigi & Mach 8 (72050) Review: 2026’s Best Set Yet? Marc Andreessen Says AI Productivity Will Trigger A Hiring Boom 3D Printing Is The Ultimate Hack To Reduce Household Spending Apple iPhone Fold: Striking Design Revealed In Leaked Photos Apple Smart Glasses: New Leak Reveals A Major Design Twist To Beat Meta Tested: The AI Coming To The Rivian R2 Quordle Hints Today: Monday, April 13 Clues And Answers
Stop Cleaning Your Data. Start Finding The Signal.
John Sviokla · 2026-04-22 · via Forbes - Innovation
Three $100 banknotes burning, amid fire, flames and ash

Cleaning data can be a waste of time: Seek out signal

getty

The real AI advantage isn’t a pristine data lake — it’s knowing which information actually changes your decisions.

The most expensive piece of advice in enterprise technology right now is five words long: get your data ready first.

It is everywhere. The World Economic Forum reports that 72 percent of enterprises plan to prioritize data foundations and pipelines as their fastest-growing AI investment this year. Gartner predicts that through 2026, organizations will abandon 60 percent of AI projects unsupported by "AI-ready" data. Cloudera’s latest global survey found that 96 percent of IT leaders report AI integration — but nearly 80 percent say their initiatives are constrained by limited data access, and only 18 percent describe their data as fully governed. A Fivetran benchmark of more than 500 senior data and technology leaders found that 73 percent of enterprise data initiatives fail to meet expectations — despite average annual data spending of $29.3 million per organization.

The diagnosis is always the same: more governance, more cleaning, more pipeline engineering. Get the data house in order, then deploy AI.

This sounds prudent. It is destroying value at scale.

The logic has a seductive surface: bad data in, bad decisions out. Nobody disagrees. But the conclusion most enterprises are drawing — that data must be cleaned, standardized, and governed before AI can be useful — inverts the actual sequence of value creation. It assumes you know which data matters before you have asked which decisions matter. And it treats AI as something that consumes clean data, rather than what it actually is: the most powerful tool ever built for finding structure in unstructured information.

The result is an enterprise data strategy that spends tens of billions of dollars a year polishing datasets that may contain no decision-relevant signal at all — while ignoring messy, unstructured sources that are rich with signal but don’t fit the governance framework. The same dynamic recently played out in dramatic fashion when OpenAI shut down Sora — a product with massive compute but no proprietary signal moat beneath it. What killed Sora at the product level is killing enterprise AI initiatives at the portfolio level. Roughly 80 percent of data lake initiatives eventually fail, degenerating into what practitioners bluntly call "data swamps." The lakes are clean. The signal was never in them.

MORE FOR YOU

There is a better question to ask before any of this spending begins. It comes from decision theory, and it has a name: the Expected Value of Perfect Information — EVPI. The framework is simple: if you had perfect information, would it materially change the decision you are about to make? If the answer is no, the information has no economic value no matter how clean it is. If the answer is yes, even messy, unstructured, "dirty" data that points toward that answer is worth more than a perfectly governed dataset that tells you nothing new.

The enterprises that apply this lens first — signal before cleaning, decisions before infrastructure — are building durable competitive advantages. The ones that don't are building expensive filing cabinets.

The Koch Principle: Why Dirty Data Can Be More Valuable Than Clean Data

Koch Industries’ Pine Bend refinery in Minnesota offers an instructive analogy. Family members who feuded over the business openly referred to their Canadian feedstock as "garbage crudes" — sulfur-laden, low-grade material that other refineries passed over. The crude was extraordinarily cheap precisely because of its high sulfur content, and few refineries could process it — but Koch sold its refined products into markets where supply was tight and prices were high. The willingness to process what others wouldn't, and the expertise to extract premium value from difficult feedstock, became one of the most durable competitive advantages in American industry.

The signal principle in AI works exactly the same way. Unstructured customer feedback, call center transcripts, satellite imagery, sensor telemetry, proprietary operational logs — these are the sulfur-laden crudes of the data world. Abundant, cheap, and largely ignored by organizations that are busy cleaning what they already have. The firms that develop the capability to extract decision-relevant signal from this feedstock will outperform those who are still polishing their master data management programs.

And here is the irony the "data readiness" consensus misses entirely: AI is the refinery. It is the tool purpose-built to find structure in unstructured data, to surface patterns in noise, to extract value from feedstock that traditional analytics cannot process. Delaying AI deployment until the data is clean is like telling Koch to stop refining until someone else removes the sulfur from the crude. The sulfur is the opportunity. The refinery is the capability. You don't sequence the cleaning before the processing. You build the processing capability and let it tell you what's worth cleaning.

What Does a Signal-First Data Strategy Look Like?

The enterprise-scale proof of this principle predates generative AI by decades. Verisk Analytics didn’t build a data lake and ask what could be done with it. The company started with the decisions the insurance industry needed to make — how to price risk, detect fraud, model catastrophe exposure, assess claims — and then systematically acquired every organization that held signal relevant to those decisions.

AIR Worldwide gave them probabilistic catastrophe signal. Xactware, embedded in 80 percent of major property insurers’ claims workflows, gave them granular repair cost signal. Jornaya gave them real-time consumer intent signal — knowing who is actively shopping for coverage before they have applied. Each acquisition answered the same question: what information, if known, would actually change an underwriter's, adjuster's, or executive's decision? That is EVPI applied as an M&A strategy.

The network effects reinforced the moat. Participating insurers contribute their own loss data to Verisk’s statistical database in exchange for access to the aggregated industry dataset. Nobody is cleaning data for its own sake. They are contributing signal to get better signal back. Governance follows the value. It always has.

Verisk's market capitalization reached over $38 billion at its peak — roughly 13 times its 2009 IPO valuation. The asset being valued is not a data lake. It is the accumulation of decision-relevant signal — and the workflows through which that signal reaches the people who need it.

What Should the Chief Data Officer Actually Do?

This reordering has profound implications for how enterprises should think about data leadership. The Chief Data Officer role has been defined, almost universally, as a governance and infrastructure function: clean the data, catalog the data, build the pipelines, manage compliance. These are real and necessary tasks. But they are second-order tasks.

The first-order task is signal identification: working backward from the decisions the enterprise needs to make and asking — with EVPI rigor — which information, if known, would actually shift those decisions and by how much. That is where competitive advantage lives. Governance and quality standards should follow signal priorities, not precede them.

Traders have always understood this. A fixed-income desk doesn’t ask whether all the data is clean before looking for yield signals. They ask what moves the price. Enterprise data leaders need to develop that same instinct — and build organizations that reward signal discovery, not just data hygiene.

Clean data, in the absence of signal, is just an expensive filing cabinet.