惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

雷峰网
雷峰网
L
LangChain Blog
GbyAI
GbyAI
F
Fortinet All Blogs
腾讯CDC
Last Week in AI
Last Week in AI
A
About on SuperTechFans
J
Java Code Geeks
Microsoft Azure Blog
Microsoft Azure Blog
博客园 - Franky
B
Blog
D
Docker
G
Google Developers Blog
月光博客
月光博客
博客园 - 三生石上(FineUI控件)
S
SegmentFault 最新的问题
Apple Machine Learning Research
Apple Machine Learning Research
酷 壳 – CoolShell
酷 壳 – CoolShell
freeCodeCamp Programming Tutorials: Python, JavaScript, Git & More
T
Tailwind CSS Blog
宝玉的分享
宝玉的分享
U
Unit 42
Blog — PlanetScale
Blog — PlanetScale
B
Blog RSS Feed

TechWire Asia

Microsoft ERP Implementation Partners Roundup: Which Is Best For Your Business? Semiconductor talent retention now comes with a price tag, and it keeps climbing AI token usage is not productivity. Here's what to measure instead Oracle brings governed AI agents to Fusion Applications Nvidia expands Japan AI infrastructure and robotics push AI Appreciation Day 2026 puts trust and governance in focus NVIDIA pours its full stack into Japan. The flip side of its China lockout? Malaysia's digital regulations are becoming a real cost for its startups Malaysia's AI data center vision: How EdgeConneX is building for the future Southeast Asia tech funding doubled to $7.4 billion. One company took most of it SK Hynix's Nasdaq listing raises $26.5 billion to fund Korea's AI memory expansion OpenAI launches GPT-5.6 for coding, cyber and science Meta rolls out Muse Image AI model for Instagram, WhatsApp, and advertisers Malaysia businesses face AI and password cybersecurity risks How AI workloads will test APAC mobile networks Enterprise AI costs don't have to spiral, argues ManageEngine Microsoft launches $2.5B Frontier Company for enterprise AI FIFA World Cup: How To Win Fans in APAC With Technology Kanga enters a new phase of global growth and launches Kanga Global Vertiv ramps up manufacturing in Johor's tightening data centre market U Mobile completes migration to own ULTRA5G network after DNB exit Anthropic Claude models launch in Microsoft Foundry on Azure Asia built the AI infrastructure boom. The BIS just flagged who's exposed if it stalls. Why Apple is lobbying Washington to buy China’s memory chips Nvidia-backed Firmus plans 170,000-GPU Batam AI data centre Taiwan robot makers march into humanoid systems IBM claims world’s first sub-1 nm chip technology using nanostack design Can Alibaba bridge Malaysia’s SME talent gap via agentic AI for business? Huawei’s new tech explains why mobile AI network tech is no longer optional Apple-Intel chip deal faces years-long production timeline
Microsoft to deploy AMD Helios AI systems on Azure
Muhammad Zulhusni · 2026-07-21 · via TechWire Asia
  • Microsoft will deploy AMD Helios on Azure.
  • Azure will add AMD EPYC VMs and Pensando networking.

Microsoft plans to deploy AMD’s Helios rack-scale computing platform on Azure under an expanded infrastructure partnership covering processors, graphics chips, networking hardware, and software. The systems will support AI inference for Microsoft, Azure customers, and Azure AI services.

Microsoft also plans to offer the Helios infrastructure through its upcoming ND MI455X v7 virtual machines. The instances are designed for production-scale reasoning, search, and agentic inference workloads.

Helios targets large-scale AI inference

Helios combines AMD Instinct MI455X GPUs, sixth-generation EPYC “Venice” processors, Pensando networking technology, and the ROCm software platform. The reference design includes 72 MI455X GPUs, EPYC host processors, and Pensando “Vulcano” networking hardware in a double-wide rack.

AMD developed the system around Meta’s Open Rack Wide specification, which was submitted through the Open Compute Project. The design provides a common rack architecture for high-density computing, power delivery, and liquid cooling.

“AMD and Microsoft have spent years building high-performance infrastructure together, and today we’re extending that partnership across the full stack of AMD AI solutions on Azure,” AMD Chair and CEO Lisa Su said. “Microsoft’s new AMD deployments mark an important milestone as we deliver leadership compute solutions to Azure customers and scale the next generation of AI infrastructure together,” Su said.

Each MI455X GPU includes 432GB of HBM4 memory and provides memory bandwidth of up to 19.6TB per second, according to AMD. A full Helios rack provides up to 31TB of combined HBM4 memory across its 72 accelerators.

AMD rates the system at up to 1.4 exaflops of FP8 compute and 2.9 exaflops of FP4 compute. The figures represent peak manufacturer specifications rather than measured Azure application performance.

The MI455X GPUs handle the primary AI calculations, while the EPYC processors support host computing, workload coordination, and data movement. Microsoft will initially use Helios for frontier-model inference across its own services, Azure AI services, and customer applications.

Although Helios supports both model training and inference, Microsoft’s announced ND MI455X v7 deployment focuses on running trained models at scale. The company has not provided a timetable for customer access to the instances.

“Customers are looking for AI infrastructure that is optimised for a wide range of workloads, from training and inference to data preparation, search, and reinforcement learning,” Microsoft Chairman and CEO Satya Nadella said. “Through our collaboration with AMD, we are expanding the Azure infrastructure portfolio with AMD Helios to give customers the performance, scale, and choice they need to build and run the next generation of AI applications,” Nadella said.

Enterprise customers will also be able to access AMD-based infrastructure through Microsoft Foundry Managed Compute. The service hosts open-source models on dedicated GPU capacity, with Microsoft managing the GPU topology, runtime, container image, and security patching.

Customers select the model, accelerator family, deployment template, and scaling settings. Managed Compute remains in public preview, has no service-level agreement, and is not currently recommended by Microsoft for production workloads.

Azure expands CPU and networking infrastructure

The partnership also covers two Azure virtual machine series powered by AMD’s sixth-generation EPYC “Venice” processors. Azure HDv2 will include nearly 500 physical EPYC CPU cores, 4TB of RAM, 32TB of local NVMe storage, and 400Gb Azure Boost networking.

Microsoft is positioning HDv2 for CPU-intensive AI workloads, including data preparation, search, reinforcement learning, agent coordination, and data pipelines. The company said the VMs will handle processes that prepare data, coordinate workloads, and support GPU-based training and inference.

Azure HXv2 will include 176 sixth-generation EPYC cores running at more than 5GHz, with 50% more addressable cache per core than the previous HX generation. Microsoft plans to offer configurations with nearly 2TB or 4TB of RAM and 800Gb InfiniBand connectivity.

HXv2 will retain AMD’s 3D V-Cache technology, which is also used in the current Azure HX series. Microsoft and AMD introduced the first HX virtual machines in 2023 for electronic design automation workloads.

The HXv2 series will target electronic design automation, scientific simulation, engineering analysis, and distributed-memory computing. Microsoft also identified register-transfer level simulation as a target workload for the new series.

The 800Gb InfiniBand connection is intended to support large-scale Message Passing Interface simulations across distributed computing environments. AMD also uses Azure HX infrastructure for electronic design automation as it develops future EPYC processors and Instinct accelerators, according to AMD Executive Vice-President and CTO Mark Papermaster.

Microsoft has not announced pricing, launch dates, or the Azure regions where HDv2 and HXv2 will initially be available.

AMD and Microsoft are also expanding their work on cloud networking. Azure already uses AMD Pensando data processing units, or DPUs, to handle infrastructure functions separately from a server’s main processors.

Pensando DPUs offload networking, storage, security, and encryption services from host CPUs. Azure Boost also moves selected networking, storage, and virtualisation processing onto dedicated hardware and software.

The expanded agreement will extend AMD’s role in Azure connection processing and backend networking, although the companies have not disclosed how each component will be deployed. Within Helios, UALink-over-Ethernet provides scale-up connectivity among the rack’s GPUs, while Pensando Ethernet components support scale-out networking between racks and clusters.

ROCm provides the software environment used to develop and run workloads on AMD accelerators. Its inclusion gives Helios a common software layer across the rack’s GPU infrastructure.

AMD expects Helios-based systems to enter volume deployment in the second half of 2026. Microsoft has not announced when ND MI455X v7 instances will become available to Azure customers or which regions will receive them first.

Want to learn more about AI and big data from industry leaders? Check out AI & Big Data Expo taking place in Amsterdam, California, and London. The comprehensive event is part of TechEx and is co-located with other leading technology events, click here for more information.

Tech Wire Asia is powered by TechForge Media. Explore other upcoming enterprise technology events and webinars here.