惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

让小产品的独立变现更简单 - ezindie.com
让小产品的独立变现更简单 - ezindie.com
V
V2EX
小众软件
小众软件
MongoDB | Blog
MongoDB | Blog
Jina AI
Jina AI
G
Google Developers Blog
H
Help Net Security
Microsoft Azure Blog
Microsoft Azure Blog
月光博客
月光博客
The GitHub Blog
The GitHub Blog
Y
Y Combinator Blog
爱范儿
爱范儿
B
Blog
云风的 BLOG
云风的 BLOG
H
Hackread – Cybersecurity News, Data Breaches, AI and More
GbyAI
GbyAI
博客园 - 叶小钗
aimingoo的专栏
aimingoo的专栏
Blog — PlanetScale
Blog — PlanetScale
奇客Solidot–传递最新科技情报
奇客Solidot–传递最新科技情报
有赞技术团队
有赞技术团队
博客园_首页
Google DeepMind News
Google DeepMind News
M
MIT News - Artificial intelligence

The Register

Shadow IT has given way to shadow AI. Enter AI-BOMs Zed team releases version 1.0 of Rust-built editor: Traditional editor and AI tool Microsoft boss tells investors the company is working to 'win back fans' What type of 'C2 on a sleep cycle' do they leave behind? Novel Chinese spy group found in critical networks in Poland, Asia NASA boss: Make Pluto A Planet Again GitHub says sorry and vows to do better as uptime slips and devs complain Age checks could turn internet into an ID checkpoint, complains Proton CEO Microsoft gives your Word documents an AI co-author you didn’t ask for Datadog digs down into GPU efficiency as AI costs soar If malware via monitor cables is a matter of national security, this might be the gadget for you Thunderbird in hand worth 2 Outlooks as fresh FOSS fave and Firefox arrive Grafana offers AI assistant for free, warns users not to go mad Right to repair champ Framework punts modular 13in laptop with Core Ultra Series 3 France's 'Secure' ID agency probes breach as crooks claim 19M records Scotland Yard can keep using live facial recognition on Londoners, say judges UK tribunal sends £2B claim accusing Microsoft of overcharging for licensing to trial Nation-states want to cause harm, not just steal cash - stop handing your cyber defenses to the cheapest contractor Murder, she wrote: Ex-FBI chief wants some ransomware crims charged with homicide Phone-to-satellite use goes into orbit, growing 25% in 8 months macOS ClickFix attacks deliver AppleScript stealers to snarf credentials, wallets Anthropic bakes memory fixes into Bun 1.1.13 as developers complain of leaks The spaghettified DBMS chart that shows Oracle's crown is slowly slipping Yet another ex-ransomware negotiator admits turning rogue after payoff from crimelords FAA grounds Blue Origin's New Glenn as it probes missed satellite delivery 'mishap' AMD's Ryzen 9 9950X3D2 Dual Edition tested: Gratuitous overkill with a price to match AI-assisted intruders pwned Vercel via OAuth abuse and a pilfered employee account Crook claims to leak 'video surveillance footage' of companies Met police trials snoop tech platform in push to cuff more London shoplifters England's school phone ban gets teeth, just in time to bite no one Adaptavist Group breach spawns imposter emails as ransomware crew claims mega-haul
ZTE builds a TCO-optimal AI factory to fuel token economy
ZTE · 2026-06-25 · via The Register

PARTNER CONTENT: Leveraging OEX architecture SuperPODs and multi-dimensional co-design to maximize tokens per second and lower total cost of ownership for scaled inference

ZTE showcased at MWC Shanghai 2026 its comprehensively boost to TPS. Powered by multiple dimensional co-design, deep optimization, and acceleration - from chips, servers, clusters, and AIDC, to software algorithms and scheduling platforms - this innovation empowers customers to build TCO-Optimal AI factories providing robust support for the efficient development of the Token Economy.

As large models enter the phase of scaled inference deployment, the "cost per Token" has emerged as the ultimate metric for measuring the commercial value of AI. ZTE proposes that a leap in Token generation efficiency can only be achieved through architectural-level innovation and system-level synergy. To that end, the OEX (Orthogonal Electrical eXchange) architecture based SuperPOD showcased at MWC Shanghai 2026 represents a milestone innovation designed to shatter computing power bottlenecks and maximize energy efficiency.

Pioneering the OEX architecture to define the next-generation super-node standard

ZTE pioneered the Orthogonal Architecture SuperPOD concept. Its OEX architecture features a midplane-free and zero-cable design to achieve physical decoupling and flexible replacement of core components such as GPUs, CPUs, and switch chips. By supporting mainstream high-speed interconnect protocols like CLink and SUE, it truly realizes "multi-chip synergy, open compatibility, and on-demand optimization". Compared with traditional architectures, OEX-based SuperPOD communication paths are shorter with lower signal loss, significantly improving overall interconnection efficiency, minimizing latency, and enhancing system reliability.

ZTE's SuperPOD single rack achieves industry-leading ultra-high-density integration of 128 GPUs, and supports scale up to 16,000 GPUs to build an extra-large-scale cluster. It meets AI training and inference requirements ranging from thousand card to ten thousand card scale, providing a solid foundation for long context, high concurrency agent scenarios.

Multi-dimensional optimization to achieve comprehensive improvements in inference efficiency and computing energy efficiency

Hardware-software synergy unleashes ultimate efficiency per watt. By leveraging a PD disaggregation and integrating technologies such as network efficiency optimization, operator optimization, and multi level KV cache, performance bottlenecks are overcome and throughput is significantly increased. In close collaboration with multiple manufacturers, heterogeneous mixed inference and system level optimization are advanced on domestic chip platforms, resulting in a comprehensive enhancement of inference efficiency and a notable increase in TPS.

Compute-storage-network synergy builds a large scale inference pool. ZTE offers the Full Series AI Server supporting high density deployment with 8 or 16 GPUs per server and 64 or 128 GPUs per rack, adapting to diverse scenarios. The AI-native KV cache is implemented through DPU hardware acceleration that enables direct GPU access to storage, achieving zero-copy data transfer, microsecond-level latency, and PB-scale scalability. Combined with intelligent prefetching and dynamic eviction mechanisms, the cache delivers a hit rate exceeding 70%, significantly boosting inference efficiency.

Building an open and evolvable AI infrastructure with ecosystem partners

ZTE emphasizes that AI computing power development must balance performance, cost, and sustainable evolution. To achieve this, ZTE's OEX-based SuperPOD adopts a "Pre Integration" model. Through Pre adaptation & Pre integration, the product adaptation and turning cycle is slashed from over one year to within six months, significantly accelerating ecosystem convergence and large-scale commercial rollout.

Efficient computing power serves as the bedrock of the Token economy, and ZTE is dedicated to ensuring that every ounce of computing power translates into tangible AI productivity. This showcase fully underscores ZTE's deep technological heritage and innovative strength in intelligent computing infrastructure, while offering customers a definitive, future-proof blueprint for building highly efficient AI factories.

Contributed by ZTE.