惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

Cyber Security Advisories - MS-ISAC
Cyber Security Advisories - MS-ISAC
G
Google Developers Blog
S
SegmentFault 最新的问题
Microsoft Security Blog
Microsoft Security Blog
J
Java Code Geeks
罗磊的独立博客
H
Hackread – Cybersecurity News, Data Breaches, AI and More
量子位
P
Proofpoint News Feed
博客园 - 【当耐特】
MongoDB | Blog
MongoDB | Blog
L
LangChain Blog
F
Fortinet All Blogs
C
Check Point Blog
博客园_首页
I
InfoQ
Jina AI
Jina AI
Blog — PlanetScale
Blog — PlanetScale
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻
酷 壳 – CoolShell
酷 壳 – CoolShell
Engineering at Meta
Engineering at Meta
美团技术团队
Vercel News
Vercel News
Apple Machine Learning Research
Apple Machine Learning Research

Fortune | FORTUNE

One man can kill Bill Ackman’s $64 billion bid for Universal Music Group—and no one knows what he’ll do | Fortune Poppi’s cofounder pitched her startup on Shark Tank while 9 months pregnant and landed a $400,000 deal—now it's worth $2 billion | Fortune Teen boys are choosing AI girlfriends over real ones for 'maximum control, zero rejection'—experts say it could make them unemployable | Fortune A United American merger is by no means impossible given the president 'loves big deals' | Fortune Reed Hastings’s planned exit from $455 billion Netflix ‘had nothing to do with’ the failed deal for Warner Bros., says Ted Sarandos | Fortune Meet Joe McCann: The high-flying crypto trader held in Tanzania after sudden death of his influencer fiancée Ashly Robinson | Fortune Gen Z is carving a different path in the housing market by doing it alone | Fortune U.S. Catholic leaders criticize Trump for ‘disparaging words’ about the pope as Vatican clash risks alienating Catholic voters | Fortune China has ‘nearly erased’ America’s lead in AI—and the flow of tech experts moving to the U.S. is slowing to a trickle, Stanford report says | Fortune Self-made millionaire behind $5 billion Skims Emma Grede says it all began with a cold call to Kris Jenner: Emma Grede—the self-made millionaire behind the $5 billion Skims empire—says it all began with an audacious cold call to Kris Jenner: ‘The difference between me and someone else is, I made it happen’ | Fortune Americans have never been this gloomy about the economy. Wall Street has never cashed in harder | Fortune ‘The college grading system [is] almost meaningless’: People see the Ivy League as an easy A and with flawed admissions standards | Fortune The CEO of $8.5 billion Japanese car giant Nissan plays the drums in a band and hits the tennis courts to destress from the top job | Fortune New York governor's take on a millionaires tax: fancy pied-à-terre second apartments worth over $5 million | Fortune Pope Leo XIV: A ‘handful of tyrants’ are ravaging earth with war and exploitation | Fortune Trump has no plan to cut the $39 trillion national debt, but he does want to cut childcare. His budget director is scrambling to clarify | Fortune China's economy grows 5% in first quarter, surprising economists to the upside | Fortune Everyone was wondering what Trump wanted more: Warsh smoothly seated at the Fed, or for Powell to pay. We have our answer | Fortune Palantir exec: the biggest mistake retailers are making with AI? Trying to do it all with one agent | Fortune American YouTuber who calls himself a 'troll' sentenced to 6 months in Korean prison for literally dancing on wartime graves | Fortune BBC plans to cut up to 2,000 jobs to save 10% of annual budget | Fortune Canva debuts a new suite of agentic tools, as the design app quietly becomes one of the world’s most used AI services | Fortune Moody's CEO: AI has a trust problem – better models won’t fix it | Fortune Top New York surgeon: Americans have better data for choosing restaurants than surgeons. That has to change | Fortune The Iran war’s fertilizer shock is hammering American farmers, and 70% can’t afford what they need for this year’s growing season | Fortune Education experts to Mamdani: Why are you foisting AI on our kids? | Fortune This CEO pirated video games as a teen and became a hacker for the Air Force. Now he’s built a $3 billion cyber firm | Fortune Teacher, blame thyself: Yale report savages Ivy League schools for destroying American trust in higher education | Fortune Fed chair nominee Kevin Warsh is worth more than $100 million and has stakes in SpaceX and Polymarket | Fortune From wool sneakers to GPUs: Allbirds’ desperate AI pivot and 600% stock surge, explained | Fortune
Exclusive: A former Apple engineer thinks AI infrastructu...
Lily Mae Lazarus · 2026-06-25 · via Fortune | FORTUNE

For months, Kleiner Perkins partner Aditya Naganath had been mulling over his investing thesis that the next wave of AI wasn’t going to be a chatbot—it was going to be software that does the work autonomously, for hours at a time, across thousands of tasks at once. The trouble was, nobody had built the plumbing for it yet. Then he met Neil Movva.

“It felt obvious to both of us that you’re going to need a different, specific inference platform built for these long-running agents,” Naganath told Fortune.

Now, six months after Naganath and Movva first chatted, Movva’s startup, Sail Research, has launched from stealth with $80 million in seed and Series A funding at a $450 million valuation, Fortune learned exclusively. Kleiner Perkins led the Series A. Sequoia, Redpoint, Theory Ventures, Vine Ventures, and CRV also participated.

Sail Research wants to fix one of AI’s expensive problems. AI infrastructure was designed for quick, single exchanges—think a chatbot answering a question. But enterprises are increasingly deploying AI agents that run autonomously for hours, reading entire codebases, screening hundreds of job candidates, or researching complex topics without a human in the loop. At that scale, enterprise AI bills have tripled even as per-token prices have fallen, because agentic workflows consume tokens at a rate 50 to 500 times higher than simple chat. Goldman Sachs forecasts a 24-fold increase in token consumption by 2030.

Movva’s solution is an end-to-end infrastructure platform built from the lowest level of the chip up. Sail writes the software that orchestrates and optimizes how AI models run on existing chips. Think of it like a highly efficient traffic system that tells the hardware exactly how to allocate its resources, squeezing far more work out of the same physical computing power.

Most AI serving platforms optimize for low latency, meaning they prioritize getting you an answer fast. Sail does the opposite, sacrificing real-time responsiveness to pack far more computing work into every unit of power. The tradeoff is deliberate: Sail can’t power a voice assistant or a live chatbot. But for agents that run for hours? Movva claims customers often seen between 3x to 10x cost improvements over comparable alternatives.

“We only care about efficiency,” Movva told Fortune. “It’s quite difficult to build an inference engine for both throughput and latency at the same time. Everyone else is optimizing for latency, and we just care about throughput.”

Movva, 28, is one of a small number of engineers who has worked at every meaningful layer of the AI stack. He watched NVIDIA pivot from gaming chips to AI silicon in 2016 and 2017. He joined Apple to work on the chip powering computer vision on a billion iPhones—then grew frustrated that Apple’s ambition topped out at animoji (the animated characters users can apply on FaceTime). From there, he went to Together AI, one of the leading open-source model inference providers, to get back to GPU-level work. What he saw there crystallized Sail’s thesis: Together had been built for interactive applications and had made every architectural trade-off accordingly. Long-horizon agents needed something built from scratch with different priorities.

Co-founder and CTO Samir Menon also comes from Apple, where he worked in security engineering at scale. The two met on the first day of freshman year at Stanford—they took the same classes, and saw the same academic counselor. Movva jokes that Menon got slightly better grades. They reunited in late 2025 to rebuild the inference stack from scratch.

Sail launched its inference service in March and has already ramped to processing trillions of tokens per week. One early customer, Detail.dev, uses Sail to run code-review agents that spend three to four hours—sometimes longer—digging through an entire codebase hunting for bugs that five-minute reviews miss. “The abundance of tokens that we provide lets them be maximally ambitious in how they scan through code bases,” Movva said.

But the competitive risk is real. Together AI is a formidable incumbent, and it’s also a Kleiner Perkins portfolio company. Naganath’s view is that the two are not in conflict: Together owns the interactive, chat-based market; Sail owns the long-running agent workload. “Being specific and purpose-built should win out in the long run,” he said. The larger threat may come from the frontier labs—Anthropic, OpenAI, and Google—which are building their own inference infrastructure and could, in theory, commoditize the layer Sail is betting on. 

Movva’s counter: token prices have been flat or rising for six months, demand for compute is growing faster than supply, and the world needs someone focused obsessively on squeezing the most intelligence out of every available GPU. “We feel an emotional pain when we see a GPU be idle or wasted in any way,” he said.

Naganath’s bull case is simple: “The belief that inference is going to be a 10x—even 100x—bigger market than it is today.”