惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

cs.CL updates on arXiv.org
cs.CL updates on arXiv.org
博客园_首页
GbyAI
GbyAI
罗磊的独立博客
Y
Y Combinator Blog
宝玉的分享
宝玉的分享
人人都是产品经理
人人都是产品经理
U
Unit 42
V
Visual Studio Blog
F
Fortinet All Blogs
小众软件
小众软件
Cyber Security Advisories - MS-ISAC
Cyber Security Advisories - MS-ISAC
L
LangChain Blog
让小产品的独立变现更简单 - ezindie.com
让小产品的独立变现更简单 - ezindie.com
Engineering at Meta
Engineering at Meta
aimingoo的专栏
aimingoo的专栏
The Cloudflare Blog
T
Tor Project blog
Martin Fowler
Martin Fowler
K
Kaspersky official blog
Scott Helme
Scott Helme
cs.AI updates on arXiv.org
cs.AI updates on arXiv.org
D
DataBreaches.Net
博客园 - Franky
阮一峰的网络日志
阮一峰的网络日志
博客园 - 【当耐特】
P
Proofpoint News Feed
N
Netflix TechBlog - Medium
美团技术团队
S
Secure Thoughts
C
Cisco Blogs
M
MIT News - Artificial intelligence
L
Lohrmann on Cybersecurity
T
Tenable Blog
N
News and Events Feed by Topic
Threat Intelligence Blog | Flashpoint
Threat Intelligence Blog | Flashpoint
C
Check Point Blog
C
Cyber Attacks, Cyber Crime and Cyber Security
Spread Privacy
Spread Privacy
S
Security @ Cisco Blogs
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻
Microsoft Security Blog
Microsoft Security Blog
A
Arctic Wolf
Hacker News - Newest:
Hacker News - Newest: "LLM"
H
Hacker News: Front Page
T
Threat Research - Cisco Blogs
Simon Willison's Weblog
Simon Willison's Weblog
cs.CV updates on arXiv.org
cs.CV updates on arXiv.org
O
OpenAI News
V
Vulnerabilities – Threatpost

The Register - Special Features

Troops’ phones gave away location data to foreign adversaries Qualcomm picks bad time to pitch a $300 laptop platform AI agents get their own phone directory built atop DNS Carnival confirms ShinyHunters cruised off with 6M customer records after April breach Google engineer accused of turning Year in Search secrets into Polymarket payday Are we human? India's cyber agency sets clock at 12 hours to tackle exploited bugs as AI turns up the heat Broadcom gets early start on WiFi 8 with next-gen wireless routing kit Are we human? Microsoft Excel champ proves he still has the formula Anthropic co-founder hallucinates ghost in the machine Anthropic co-founder hallucinates ghost in the machine NASA plans Moon Base buildout with rovers, drones, cargo landers MyPillow must decide whether to be firm or soft as ransomware crims demand pay Starship shows it can deploy satellites, but Moon mission clock still ticks Huawei's chip law looks less like Moore and more like marketing Experts pour cold borscht on Farage's Russian hack claim Logitech unveils a cushioned mouse for all-day use AI eyes scanning for bugs create a worrisome Linux security trend A Russian speaker and jailbroken Gemini went on a hacking spree and emptied at least one MAGA victim's crypto wallets AI datacenter boom collides with US grid reality Media giant settles for $930k amid user-snooping allegations AT&T sues to ditch Cali copper phone lines to save billions FBI warns of Kali365 as device code phishing soars Techie claims Trump Mobile website was leaking thousands of people's data BOFH: Vibe-coded solutions arrive for problems nobody has Dems slam Trump for making cybersecurity hold out the tin cup while splurging on ballroom and Jan. 6 'slush fund' Google explains how it will infuse ads into AI answers AI is getting pricey, but relief is coming, but not for you Deus ex machina: Half of US Christians trust AI's spiritual advice Attackers spill plaintext passwords of 46k Myspace93 users after 2021 breach Apple adds AI smarts to Voice Control, VoiceOver and Magnifier ahead of Accessibility Day Microsoft open-sources agentic AI safety tools OpenAI wants upfront cash for guaranteed AI capacity Fedora: Microsoft is all aboard, but Deepin is dumped Bye-bye, Gemini CLI; Google nudges devs toward Antigravity Plex appeal fades as Lifetime Pass jumps to $750 AI sackings reach New Zealand, which will use it to eject 14 percent of government staff Anthropic’s Stainless steal tightens grip on AI dev tooling Are we human? Google touts tokenmaxxing, huge capex, and AI agents at I/O America's top cyber-defense agency left a GitHub repo open with with passwords, keys, tokens – and incredibly obvious filenames America's top cyber-defense agency left a GitHub repo open with passwords, keys, tokens – and incredibly obvious filenames Shadow AI invades the workplace, up 4x in the last year Microsoft refreshes Surface for Business lineup, starts AI PC upsell at $1,499 Broadcom finds a VMware customer willing to stick around: London Stock Exchange 468k records allegedly stolen from Portugal’s postal carrier Baidu says the quiet part out loud – you can’t build AI infrastructure, so clouds can cash in Shai-Hulud copycat worm infects yet another npm package Uncle Sam's next big super might not use GPUs Are we human? Datacenters slurping up so much juice they boosted prices 75% in largest US energy market MPs want social media treated more like unsafe toys than harmless apps Cerebras’ wafer-scale AI bet delivers blockbuster IPO Nobody believes the 'criminals and scumbags' who hacked Canvas really deleted stolen student data Anthropic tosses agents into the API billing pool Jen Easterly, cybersecurity's 'relentless optimist,' hopes feds come back to RSAC next year Jen Easterly, cybersecurity's 'relentless optimist' Smooth criminals talking their way into cloud environments, Google says Voice phishing skyrockets as smooth crims talk their way in RSAC 2026: Uncle Sam backs out, AI agents everywhere RSAC 2026: Uncle Sam backs out, AI agents everywhere Decoding Nvidia's Groq-powered LPX and the rest of its new rack systems A closer look at Nvidia's Groq-powered LPX rack systems Nvidia slaps $20B Groq tech into massive new LPX racks to speed AI response time Nvidia slaps Groq into new LPX racks for faster AI response AI Burning Man happens next week – what to expect at Nvidia GTC 2026 Nvidia GTC 2026: What to expect at AI Burning Man Unaccounted-for AI agents are being handed wide access Unaccounted-for AI agents are being handed wide access Google to foist Gemini pane on Chrome users Google to foist Gemini pane on Chrome users Yes, you can build an AI agent – here's how, using LangFlow How to build an AI agent using LangFlow Clawdbot becomes Moltbot, but can’t shed security concerns Clawdbot becomes Moltbot, but can’t shed security concerns Gartner questions if Salesforce AI will stay all-you-can-eat Gartner questions if Salesforce AI will stay all-you-can-eat Claude supports MCP Apps, presents UI within chat window Claude supports MCP Apps, presents UI within chat window Cursor is better at marketing than coding Cursor is better at marketing than coding Feds skipping infosec industry's biggest conference, RSAC AI is rewriting how power flows through the datacenter All aglow about DCs, investors launch $300M at microreactor startup Radiant bags $300M-plus to commercialize its microreactors Why do bit barns keep bumping up our bills, Senators ask DC operators Senate trio questions DC operators over rising energy costs Building the AI factory datacenter Delays? What delays? Oracle insists its $300B cloud contract with OpenAI is on track Oracle insists its $300B contract with OpenAI is on schedule Salesforce willing to lose money on AI to lock in customers Salesforce willing to lose money on AI to lock in customers Galactic Brain space datacenter coming in 2027, pledges startup Aetherflux Galactic Brain space datacenter promised in 2027 Activist groups urge Congress to pause datacenter buildouts Activist groups urge Congress to pause datacenter buildouts Bezos-backed Unconventional AI addresses datacenter power Bezos-backed Unconventional AI addresses datacenter power AWS re:Invent keynote: Matt Garman bores, then thrills
ZTE builds a TCO-optimal AI factory to fuel token economy
ZTE · 2026-06-25 · via The Register - Special Features

PARTNER CONTENT: Leveraging OEX architecture SuperPODs and multi-dimensional co-design to maximize tokens per second and lower total cost of ownership for scaled inference

ZTE showcased at MWC Shanghai 2026 its comprehensively boost to TPS. Powered by multiple dimensional co-design, deep optimization, and acceleration - from chips, servers, clusters, and AIDC, to software algorithms and scheduling platforms - this innovation empowers customers to build TCO-Optimal AI factories providing robust support for the efficient development of the Token Economy.

As large models enter the phase of scaled inference deployment, the "cost per Token" has emerged as the ultimate metric for measuring the commercial value of AI. ZTE proposes that a leap in Token generation efficiency can only be achieved through architectural-level innovation and system-level synergy. To that end, the OEX (Orthogonal Electrical eXchange) architecture based SuperPOD showcased at MWC Shanghai 2026 represents a milestone innovation designed to shatter computing power bottlenecks and maximize energy efficiency.

Pioneering the OEX architecture to define the next-generation super-node standard

ZTE pioneered the Orthogonal Architecture SuperPOD concept. Its OEX architecture features a midplane-free and zero-cable design to achieve physical decoupling and flexible replacement of core components such as GPUs, CPUs, and switch chips. By supporting mainstream high-speed interconnect protocols like CLink and SUE, it truly realizes "multi-chip synergy, open compatibility, and on-demand optimization". Compared with traditional architectures, OEX-based SuperPOD communication paths are shorter with lower signal loss, significantly improving overall interconnection efficiency, minimizing latency, and enhancing system reliability.

ZTE's SuperPOD single rack achieves industry-leading ultra-high-density integration of 128 GPUs, and supports scale up to 16,000 GPUs to build an extra-large-scale cluster. It meets AI training and inference requirements ranging from thousand card to ten thousand card scale, providing a solid foundation for long context, high concurrency agent scenarios.

Multi-dimensional optimization to achieve comprehensive improvements in inference efficiency and computing energy efficiency

Hardware-software synergy unleashes ultimate efficiency per watt. By leveraging a PD disaggregation and integrating technologies such as network efficiency optimization, operator optimization, and multi level KV cache, performance bottlenecks are overcome and throughput is significantly increased. In close collaboration with multiple manufacturers, heterogeneous mixed inference and system level optimization are advanced on domestic chip platforms, resulting in a comprehensive enhancement of inference efficiency and a notable increase in TPS.

Compute-storage-network synergy builds a large scale inference pool. ZTE offers the Full Series AI Server supporting high density deployment with 8 or 16 GPUs per server and 64 or 128 GPUs per rack, adapting to diverse scenarios. The AI-native KV cache is implemented through DPU hardware acceleration that enables direct GPU access to storage, achieving zero-copy data transfer, microsecond-level latency, and PB-scale scalability. Combined with intelligent prefetching and dynamic eviction mechanisms, the cache delivers a hit rate exceeding 70%, significantly boosting inference efficiency.

Building an open and evolvable AI infrastructure with ecosystem partners

ZTE emphasizes that AI computing power development must balance performance, cost, and sustainable evolution. To achieve this, ZTE's OEX-based SuperPOD adopts a "Pre Integration" model. Through Pre adaptation & Pre integration, the product adaptation and turning cycle is slashed from over one year to within six months, significantly accelerating ecosystem convergence and large-scale commercial rollout.

Efficient computing power serves as the bedrock of the Token economy, and ZTE is dedicated to ensuring that every ounce of computing power translates into tangible AI productivity. This showcase fully underscores ZTE's deep technological heritage and innovative strength in intelligent computing infrastructure, while offering customers a definitive, future-proof blueprint for building highly efficient AI factories.

Contributed by ZTE.