惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

博客园 - 聂微东
宝玉的分享
宝玉的分享
Apple Machine Learning Research
Apple Machine Learning Research
罗磊的独立博客
Last Week in AI
Last Week in AI
WordPress大学
WordPress大学
博客园 - 【当耐特】
大猫的无限游戏
大猫的无限游戏
小众软件
小众软件
博客园 - 司徒正美
博客园 - Franky
爱范儿
爱范儿
钛媒体:引领未来商业与生活新知
钛媒体:引领未来商业与生活新知
freeCodeCamp Programming Tutorials: Python, JavaScript, Git & More
T
Tailwind CSS Blog
Hugging Face - Blog
Hugging Face - Blog
Jina AI
Jina AI
量子位
博客园 - 叶小钗
博客园_首页
月光博客
月光博客
博客园 - 三生石上(FineUI控件)
The Cloudflare Blog
奇客Solidot–传递最新科技情报
奇客Solidot–传递最新科技情报

SiliconANGLE

Will agentic AI governance run amok? The lesson of Asimov’s Three Laws - SiliconANGLE AI + quantum, Amazon vs. Starlink and the wide-open US-China internet battle - SiliconANGLE Team Cymru launches Total Insights Feed to replace legacy threat intelligence lists - SiliconANGLE AI Mode in Chrome adds split-screen view to enhance the web search experience - SiliconANGLE Resolve AI raises $40M at $1.5B valuation to optimize production environments - SiliconANGLE How Zscaler and OpenAI turn zero-trust security into an AI accelerator - SiliconANGLE OpenAI ratchets up Codex's agentic capabilities to rival Claude Code - SiliconANGLE Anthropic launches Claude Opus 4.7 with coding, visual reasoning improvements - SiliconANGLE Slash raises $100M at a $1.4B valuation to expand AI-powered banking platform for online businesses - SiliconANGLE Canva unveils Canva AI 2.0, recasting its platform as an agentic system for work - SiliconANGLE Data center, consumer device chips boost TSMC’s revenue - SiliconANGLE Mission-critical security cannot be bolted on, says Oracle - SiliconANGLE Agentic infrastructure reshapes enterprise AI - SiliconANGLE Data quality, and data freedom, foundational for AI success - SiliconANGLE Data trust is a bedrock in successful, scalable AI outcomes - SiliconANGLE Google introduces new agentic AI-ready tools and resources for Android developers  - SiliconANGLE Agentic AI orchestration separates winners from laggards - SiliconANGLE Data-driven tools turning the tide against human trafficking - SiliconANGLE Achieving trusted AI development goes beyond 'vibes' - SiliconANGLE Impinj boosts edge computing power in updated R700 RAIN RFID reader - SiliconANGLE Certinia powers professional services with AI - SiliconANGLE Antioch prepares to accelerate simulated testing for autonomous robots after raising $8.5M - SiliconANGLE Developer tooling startup Expo nabs $45M investment - SiliconANGLE Solidroad lands $25M to bring AI to customer support interactions - SiliconANGLE DuploCloud lands compliance and AI governance certifications as enterprise buyers tighten scrutiny - SiliconANGLE Lua lands $5.8M to help businesses build and manage AI agent workforces - SiliconANGLE Best of frenemies: Oracle's and AWS' clouds unite with dedicated, private connectivity - SiliconANGLE NIST shifts National Vulnerability Database to risk-based triage as CVE submissions hit record levels - SiliconANGLE Cisco goes to the races with new Churchill Downs multiyear partnership - SiliconANGLE Susecon 2026 will tackle the future of open-source platforms - SiliconANGLE
AI inference provider Baseten reportedly raising $1.5B in...
by Maria Deutscher · 2026-06-19 · via SiliconANGLE

AI inference provider Baseten reportedly raising $1.5B in funding

Baseten Inc., a startup with a platform for running artificial intelligence inference workloads, is raising $1.5 billion in funding.

The Wall Street Journal reported today that Altimeter Capital, Conviction, Spark Capital, Sands Capital and Wellington Management are co-leading the deal. It’s unclear whether there are additional participants. Some of the investors are buying shares at an $11 billion valuation while the other backers’ term sheets specify a $13 billion valuation.

Setting up a cloud-based inference cluster involves a significant amount of work. Developers have to provision graphics cards, configure them, link them together and install a large number of software tools. Baseten provides a platform that automates the workflow. The software is available as a managed service and as a standalone application that companies can deploy in their public cloud environments.

Baseten’s platform is powered by three core modules the company calls inference engines. They optimize the performance of customers’ AI models and collect data about technical issues.

The first inference engine, BIS-LLM, is designed power large language models with a mixture of experts architecture. A mixture of experts LLM comprises multiple neural networks that are each geared towards different tasks. BIS-LLM improves the efficiency of such models by optimizing their KV cache, a data structure that stores information necessary for inference. When a model’s token usage increases, BIS-LLM automatically provisions more hardware.

The second inference engine is called Engine-Builder-LLM. It’s optimized for dense LLMs, which are models that comprise a monolithic collection of artificial neurons rather than multiple neural networks. AI models usually generate output one token at a time. Engine-Builder-LLM uses a technology called lookahead decoding to generate multiple tokens at once, which speeds up processing.

The third core inference engine, BEI, is geared towards simpler AI models. It can power embedding models, which turn raw data into a format that LLMs understand, as well as data classification and search models.

Baseten uses a software module called MCM to spread inference workloads across multiple public clouds. If one of the clouds experiences an outage, MCM reroutes prompts to the platforms that are still online. According to Baseten, the technology’s ability to switch providers is also handy when a company’s main public cloud has a shortage of graphics cards.

The platform provides out of the box support for several dozen open-source AI models. Additionally, customers can deploy custom algorithms using a tool called Truss. It automates the task of packaging an LLM into a Baseten-compatible format.

Baseten can not only perform inference with custom LLMs but also train them. According to the company, its platform includes a backup feature that periodically saves copies of a neural network while it’s being trained. If a technical issue crops up, developers can restore the most recent backup copy instead of starting the training workflow from scratch.

The funding comes less than six months after its previous raise. The $300 million investment included contributions from Nvidia Corp. and CapitalG, Alphabet Inc.’s growth-stage startup investment arm. 

Photo: Baseten

A message from John Furrier, co-founder of SiliconANGLE:

Support our mission to keep content open and free by engaging with theCUBE community. Join theCUBE’s Alumni Trust Network, where technology leaders connect, share intelligence and create opportunities.

  • 15M+ viewers of theCUBE videos, powering conversations across AI, cloud, cybersecurity and more
  • 11.4k+ theCUBE alumni — Connect with more than 11,400 tech and business leaders shaping the future through a unique trusted-based network.

About SiliconANGLE Media

SiliconANGLE Media is a recognized leader in digital media innovation, uniting breakthrough technology, strategic insights and real-time audience engagement. As the parent company of SiliconANGLE, theCUBE Network, theCUBE Research, CUBE365, theCUBE AI and theCUBE SuperStudios — with flagship locations in Silicon Valley and the New York Stock Exchange — SiliconANGLE Media operates at the intersection of media, technology and AI.

Founded by tech visionaries John Furrier and Dave Vellante, SiliconANGLE Media has built a dynamic ecosystem of industry-leading digital media brands that reach 15+ million elite tech professionals. Our new proprietary theCUBE AI Video Cloud is breaking ground in audience interaction, leveraging theCUBEai.com neural network to help technology companies make data-driven decisions and stay at the forefront of industry conversations.