惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

aimingoo的专栏
aimingoo的专栏
S
Securelist
博客园 - Franky
Cyber Security Advisories - MS-ISAC
Cyber Security Advisories - MS-ISAC
IT之家
IT之家
GbyAI
GbyAI
Microsoft Azure Blog
Microsoft Azure Blog
The Cloudflare Blog
云风的 BLOG
云风的 BLOG
N
News and Events Feed by Topic
AI
AI
cs.AI updates on arXiv.org
cs.AI updates on arXiv.org
Schneier on Security
Schneier on Security
Attack and Defense Labs
Attack and Defense Labs
Vercel News
Vercel News
腾讯CDC
Google DeepMind News
Google DeepMind News
K
KPMG report finds enterprise disconnect between AI and its ROI | CIO
M
MIT News - Artificial intelligence
WordPress大学
WordPress大学
freeCodeCamp Programming Tutorials: Python, JavaScript, Git & More
N
Netflix TechBlog - Medium
量子位
S
Schneier on Security
Hacker News: Ask HN
Hacker News: Ask HN
Cyberwarzone
Cyberwarzone
S
Security Affairs
钛媒体:引领未来商业与生活新知
钛媒体:引领未来商业与生活新知
N
News and Events Feed by Topic
T
Tenable Blog
PCI Perspectives
PCI Perspectives
MyScale Blog
MyScale Blog
L
Lohrmann on Cybersecurity
cs.CL updates on arXiv.org
cs.CL updates on arXiv.org
C
Cyber Attacks, Cyber Crime and Cyber Security
W
WeLiveSecurity
N
News | PayPal Newsroom
P
Proofpoint News Feed
O
OpenAI News
C
CERT Recently Published Vulnerability Notes
B
Blog
Cisco Talos Blog
Cisco Talos Blog
Microsoft Security Blog
Microsoft Security Blog
V
Visual Studio Blog
MongoDB | Blog
MongoDB | Blog
大猫的无限游戏
大猫的无限游戏
A
Arctic Wolf
Y
Y Combinator Blog
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻
Spread Privacy
Spread Privacy

Climate Drift

$41B in May. Wall Street's largest clean energy IPO ever. A €2.2B carbon-neutral lithium close. And a $500M hedge fund bet on El Niño. Four weeks to a louder climate voice The Missing Unit $55B in April. America's $1B nuclear IPO. The largest reforestation fund ever closed. And a $5B green ammonia pledge in Egypt. Your voice is load-bearing infrastructure When thought leadership isn’t all about you Your visibility gap is costing the climate $22B in March. America's $450M nuclear bet on AI power. Europe's first iron fuel plant. And a gene-edited banana 75 years in the making just raised $105M. 100 Fun Facts About the Grid Ship Fast, Ship Right Off the Record - Episode #9 - AI Energy Race 💸 Follow the Money: February 2026 The Methane Gap The $49 Billion Chocolate Fix What is powering AI? Off the Record: Climate Edition – VC doesn't work for Climate Tech. Here's 3 new solutions Off the Record: Climate Edition – How to Actually Raise Your Next Round 💸 Follow the Money: January 2026 Off the Record: Climate Edition – Solar Punk in Africa What if seaweed could build its own farm? Off the Record: Climate Edition – Want a Climate job? Give Away Your Best Ideas The hard part of hard tech Off the Record: Climate Edition – AI Surprise + VC Panic at The Drop The Executive Climate Playbook: How senior talent gets deployed
The next AI infrastructure opportunity is unlocking what we already have
Ananya Chopra · 2026-07-02 · via Climate Drift

👋 Welcome to Climate Drift: your cheat-sheet to climate. Each edition breaks down real solutions, hard numbers, and career moves for operators, founders, and investors who want impact. For more: Community | Accelerator | Open Climate Firesides | Deep Dives

Hey there 👋
Skander here.

The average U.S. power grid runs at about 30% utilization. Roughly 70% of it sits idle on a normal day, even while everyone in the room swears there’s no room left to plug in one more data center.

Somewhere in that gap sits a large, cheap answer to the AI power crunch, if you can find it.

GridCARE looked at one slice of National Grid, the same network with a years-long waiting list, and found 650 megawatts already sitting there.
Phaidra ran cooling software on live NVIDIA clusters and freed up enough thermal headroom that, on a gigawatt site, raising the loop temperature ten degrees hands 67 megawatts back to compute.
Emerald AI puts the full number at 100 gigawatts of usable capacity hiding on the existing grid.

The AI buildout is the biggest new load hitting the grid, and how much of it gets served by fresh gas versus capacity we already own shapes the emissions math for the next decade. If software can surface the stranded gigawatts, the trillion-dollar pour gets smaller, and a good chunk of the panic about AI eating the grid turns out to be a measurement problem.

We’ve watched a version of this before. Telecom companies buried fibre through the dotcom boom, betting traffic would catch up. Most of it sat dark for years and looked like the biggest overbuild in history, right until streaming and cloud arrived and turned it into the backbone of the internet.

Ananya’s read today is that AI infrastructure is at its own dark fibre moment, except the buried asset this time is gigawatts, on the grid and inside the data centers we’ve already built.

Ananya is an ex-consultant, VC and operator who works at the seam of emerging technology, strategy, and company building. She wrote this while actively looking for where to build next in the space, which is about the best reason to trust the map she’s drawn.

She walks through the companies turning software into grid infrastructure, the reasons this capacity stayed stranded for so long, and the bets she’s making anyway, open questions and all.

Today we’re looking at:

  • Why a data center can be financed and built in about 18 months but takes five to seven years to power

  • The 70% of the grid that sits idle, and why the interconnection process was never built to see it

  • How ten degrees of cooling water frees 67 megawatts on a single gigawatt site

  • The four reasons the capacity stayed stranded: nobody’s incentives, justified fear, the cracks between systems, and a design built for a world where power was free

  • Why Jensen wants to retire PUE and measure tokens per watt instead

  • Where the value actually lands, and which layer is worth owning

🌊 Let’s dive in

First: Who is Ananya?

As an ex-consultant, VC and operator, Ananya Chopra has worked at the intersection of emerging technologies, strategy and company building. She’s drawn to 0 to 1 opportunities that require systems thinking, synthesis and cross-functional execution. She enjoys developing conviction through curiosity, building from ambiguity, and turning ideas into businesses. Long term, she hopes to build companies solving consequential technology and infrastructure challenges.

Earlier this year, GridCARE looked at a slice of National Grid’s network, the same one with a years-long queue of projects waiting to plug in, and found 650MW of capacity. A couple of years ago, DeepMind pointed reinforcement learning at the cooling systems in Google’s data centers - it cut cooling energy by 40% leading to a 15% improvement in total facility overhead, using sensor data that already existed inside the building. Phaidra ran cooling agents on production NVIDIA Grace Blackwell clusters with CoreWeave and Applied Digital, cut thermal-spike overshoot by 75-80% and handed the recovered power back to compute: on a 1 GW factory, lifting the water from 25°C to 35°C frees roughly 67 MW for IT, about $3.8 billion a year in additional output.

During the dotcom boom, telecom companies buried staggering amounts of fibre-optic cable into the ground, betting that internet traffic would explode. Then the bubble popped, and most of that glass just sat there, dark. For a few years it looked like one of the largest overbuilds in history. And then streaming happened. And cloud happened. And demand finally caught up to the cable that was already in the dirt, and that dark fibre quietly became the backbone of the modern internet.

I’ve come to believe AI infrastructure is sitting on its own dark fibre moment right now. Except this time, it’s gigawatts on the grid and inside the data centers we’ve already built.

Here’s the interesting part. The number a handful of companies are now putting on the table:

Emerald AI says power-flexible data centers could free up as much as 100GW on the existing U.S. grid. Hammerhead AI says it can squeeze 20–30% more compute out of the same grid allocation.

None of them are arguing against the buildout. They’re saying that alongside the trillion-dollar concrete-pour, a massive amount of usable capacity already exists and software can help unlock it.

Which raises the question: if the capacity is right there, why is nobody using it?

Start with the grid. Stanford research pegs average grid utilization at around 30%. Yes, 70% of the grid sits idle under normal conditions, managed conservatively, invisible to the interconnection process that decides what gets built. The famous 2,600-gigawatt interconnection queue, about 2x total U.S. installed capacity, gets cited as proof the grid is full.

Now walk into the data centers. Two-thirds of organizations report peak GPU utilization under 70%. Real-world GPU usage usually runs 10–40%. Most colocations idle along at 30–50%. Even the best hyperscalers hold it above 60–70%. And almost nobody publishes their real utilization numbers, so the entire industry is planning around a baseline nobody can actually see.

A lot of that “we’re out of room” is existing infrastructure that could carry load but never shows up in the analysis, because the analysis was never built to find it.

WHY “WE’RE OUT OF ROOM” IS A MEASUREMENT PROBLEM

Morgan Stanley estimates roughly $4 billion in GPU-tenant revenue unlocked for every year you accelerate time-to-power. As Jensen Huang said, “Every data center in the future will be power-limited. Your revenues are power-limited.”

When power availability was of no consequence, excess capacity could be used to engineer reliability. In the era of AI and power-constrained data centers, software can enable a way to ensure reliability while maximizing output and unlocking more revenue.

Sense a swing and react in milliseconds instead of minutes, and you can operate against the real limit instead of a conservative guess. The proof is in the response times: Karman caps power in 20 milliseconds where software-only tools take two seconds; Phaidra reacts in under ten where conventional controllers take three to five minutes; Emerald answers a grid signal in under a minute. Each gap is margin you no longer have to pad. Speed is capacity.

The capacity is stranded for rational reasons and which is exactly why it’s been so sticky.

Suspect #1: Nobody’s incentives point at the gold. A colocation provider passes power costs through to tenants at a markup. If better thermal management recovers 20% more compute, who actually wins hard enough to go re-engineer a system that currently works? The operator gets capacity to sell. The tenant gets headroom. Neither feels the pain sharply enough to force it. Blurry incentives are a fantastic preservative for inefficiency.

Suspect #2: The fear is justified. An NVL72 rack pulls 132 kilowatts. A few years ago, the average rack was 8-17 kW; a pre-GPU server rack was more like 5-10. Density more than doubled between 2021 and 2024, and AI racks now routinely blow past 50 kW. At that power density, the terror of a thermal event is competent risk management. Running cooling at max and provisioning for the worst case is what a smart operator does when the downside is fried hardware and tenants walking. Any software that asks them to ease off that baseline has to clear a high bar.

Suspect #3: The gold lives in the cracks between systems. In a typical data center, the building management system, the power distribution units, and the compute scheduler came from different vendors, on different standards, run by different teams, optimizing different things. Nobody’s job is to optimize across all three. The waste lives precisely at the seams which is exactly why has not belonged to anyone.

Suspect #4: The whole thing was designed for a different universe. Data centers were engineered for six nines uptime, worst-case provisioning at every layer. Google’s Amin Vahdat framed the tradeoff in a very interesting way: “four nines of availability and half the capacity, or two nines of availability and twice the capacity.” The design assumption that was never updated is creating gigantic costs.

And here’s the metric that lets everybody keep not-seeing it: PUE. Power Usage Effectiveness measures how efficiently a facility turns total power into IT load. It says nothing about what that IT load actually produces. Two data centers with identical PUE can crank out wildly different amounts of useful compute. Google’s fleet runs ~1.09, AWS ~1.15, Microsoft ~1.16 and the industry average is ~1.54. That gap is pure overhead, paid on every single workload, forever.

Jensen Huang says tokens per watt drives AI factory revenue. Performance  drives token cost.

At GTC 2026, Jensen Huang proposed the replacement: tokens per watt. How much useful AI work you get per unit of power. That’s the real unlock here. Software makes these tradeoffs visible and actionable across boundaries that were never designed to talk to each other. It lights the fibre.

Data centers have always run below theoretical capacity. So why is this suddenly a concern? Three shifts turned a chronic inefficiency into a five-alarm fire.

1. Training became inference. Training is the easy roommate: sustained, predictable, high-density GPU use over long batch runs. The infrastructure handles it fine. Inference is the chaotic one - bursty, latency-sensitive, spiking orders of magnitude in minutes based on what users do. And inference is taking over. By 2027, inference demand is projected to hit 400% of 2022 training levels. The cooling logic, power contracts, and schedulers we inherited from the training era are not optimized for the job they’re now being handed.

2. GPUs got easier; power got harder. The previous era was built on GPU scarcity: control the chips, control the pricing. But the binding constraint now is power. Grid connection in primary markets runs 4+ years. Transformers average 128-week lead times and in tight markets, four years. Meanwhile an AI facility can be financed and built in 12-24 months. The need of the hour is getting data centers connected on a timeline that matches how fast AI gets built.

3. Centralized training became distributed inference. Training lived in a handful of giant clusters run by hyperscaler-grade teams who could optimize everything. Inference needs to be near users (latency) and near data (sovereignty, compliance). So, the map is fragmenting into dozens of mid-sized regional facilities, 10-50MW, and as inference takes over, flexibility stops being something you do in time at one site and becomes something you do in space across many. There is no cookie cutter model and every site requires its own optimization.

Go back to Suspect #3 for a second. The gold hiding in the cracks between systems. At a first glance, the AI infrastructure looks like a stack:

  • Physical layer at the bottom: grid gates cooling, cooling gates compute; each a hard ceiling on the one above, operating in silos.

  • Software above: Orchestration, data, inference. They continuously push demand, constraints, and tradeoffs back and forth.

  • Governance all encompassing.

But it’s more like an interconnected web. A constraint at the grid eventually shows up in inference costs. A cooling breakthrough creates additional compute capacity without adding a single watt of generation. The cross-layer dependencies nobody owns are exactly where software earns its keep. It can make hidden constraints visible, tradeoffs measurable, and optimization possible, thereby increasing the amount of productive compute that can be generated from the same megawatt.

That coordination is increasingly happening inside the facility itself. Layer 2 is evolving from facility management into infrastructure orchestration. Historically, cooling, electrical systems and compute have been operated independently. As AI pushes facilities toward their physical limits, the value shifts to software that coordinates them as one system. That’s why I think this becomes one of the most interesting new infrastructure layers to emerge.

The 2,600-gigawatt queue gets read as “the grid is overwhelmed.” But today’s power constraint issues are heavily temporal and only affect the grid for a few hours of the year.

Remember the Stanford number: ~30% average utilization, ~70% idle under normal conditions. GridCARE’s platform, physics-based AI that it says evaluates “quadrillions of grid conditions in real time”, exists to drag that hidden capacity into the light. That’s how it found the 650 megawatts on National Grid.

Emerald AI is attacking the same layer by flexing data centers. Its Conductor platform orchestrates AI workloads across one or more facilities, flexing operations in real time against grid conditions while a digital-twin simulator guards performance. The mechanism, in plain terms: some workloads can wait, some can move, batteries can be called up, all choreographed at once to hand the grid a precise, verifiable flexibility response.

They proved it where it counts. Through EPRI’s DCFlex Initiative in Phoenix, alongside NVIDIA, Oracle, and the Salt River Project utility, Emerald showed a data center could ramp its power down during a grid stress event without wrecking AI workload quality. The follow-on tells you how seriously the ecosystem is taking it. In October 2025, NVIDIA, EPRI, Digital Realty, and PJM lined up with Emerald on a “Power-Flexible AI Factory” reference design and put it into steel at Aurora, a 96 MW Digital Realty build in Manassas, Virginia, opening in the first half of 2026. By CERAWeek in March 2026, the coalition had widened to the supply side: AES, Constellation, Invenergy, NextEra, Nscale, and Vistra signed on to pioneer flexible, hybrid AI factories. With NVIDIA co-designing the product and the grid’s biggest generators line up behind it, it’s a serious signal that Emerald is onto something.

Emerald AI Orchestrates AI Factories to Help Relieve Grid Stress

AI workloads were flexed during grid stress while maintaining application performance, demonstrating that software can act as grid infrastructure

Today, compute, cooling, and power are managed as three separate systems. The losses live in that gap, the compute per megawatt i.e. how much productive AI output a facility wrings from each megawatt of contracted power.

In 2026 Phaidra ran cooling agents on production NVIDIA Grace Blackwell clusters with CoreWeave and Applied Digital, cut thermal-spike overshoot by 75-80%, and the agents self-train, pre-learning on a digital twin and then tuning to the live site in hours, not years. Riding that stability, you can raise loop temperatures and hand the recovered power back to compute: on a 1 GW factory, lifting the water from 25°C to 35°C frees roughly 67 MW for IT, about $3.8 billion a year in additional output. The efficiency is already sitting in most facilities. The cross-layer visibility to act on it is what solves for it.

Two companies are building toward that integrated control layer from different doors:

Hammerhead’s ORCA platform (Orchestrated RL Control Agents) coordinates on-site generation, UPS, cooling distribution, racks, and GPUs simultaneously and claims 20–30% more compute output from the same grid allocation, now framed as token throughput rather than raw utilization.

Utilidata’s Karman attacks oversized power infra with resolution: a million samples a second, power-capping in under 20 ms versus ~2 seconds for software-only tools. Fast enough to ride the AI power spikes that slower tools never even see, which lets a facility safely run into the headroom it’s been paying to keep idle. Utilidata’s claim: ~50% more compute on the same nameplate, and every 1 MW of unused budget left on the table is north of $20M a year in stranded revenue.

Millisecond visibility allows operators to safely utilize power headroom that conventional monitoring misses

  • How much capacity can actually be unlocked and who is incentivized to unlock it?
    The recovery rate and the adoption path remain uncertain. The most compelling efficiency gains come from highly optimized environments, making it unclear how much translates to the open market with different stakeholders. My bet: the opportunity is large enough to matter, but adoption starts where incentive, control, and urgency converge.

  • Where does the value land?
    The split between vendor, operator, and customer is unsettled, and gain-share just trades the budget objection for a measurement war over the counterfactual. My bet: whoever nails trust/verification wins the layer.

  • Can a utility-facing business be venture-scale?
    The grid-facility seam means slow, regulated buyers, long cycles, and glacial procurement. But GridCARE’s $64M and Emerald’s $68M in sixteen months suggest the capital markets are starting to bet yes.

  • Does this seam stay a startup’s to own or get absorbed?
    As hyperscalers internalize it, Schneider and Vertiv bundle it, and NVIDIA folds it into the platform? My bet: the optimizer commoditizes, but the trusted cross-party verification layer stays independent. That’s the seam worth owning.

Strip everything above to its skeleton and you get one sequence: visibility → measurement → optimization → utilization → speed-to-power.

GridCARE, Emerald, PADO, and Hammerhead look like four companies solving four problems at different layers. I’m increasingly convinced they’re all responding to the same realization: significant capacity already exists inside infrastructure we’ve already built, and the bottleneck to reaching it is software.

What I find the most interesting layer about to emerge is the control plane sitting between power, cooling, and compute inside a single facility, and how it eventually stretches across a network of facilities. As inference scatters into dozens, then hundreds, of regional deployments, the problem flips from one facility run well to many facilities coordinated intelligently. Compute and energy stop looking like separate operating decisions made site by site and start looking like a single portfolio to optimize. Tokens per watt as a fleet metric, not a facility metric becomes the number that runs the whole thing. That’s Part 2.

If you’re operating infrastructure in this space, building at any of these layers, or thinking about where to allocate capital, I’d love to compare notes. I’m also keen to build in the space and am actively exploring operating roles. If you’re working on a problem where you think I could be useful, I’d love to connect.

Discussion about this post

Ready for more?