惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

W
WeLiveSecurity
Jina AI
Jina AI
博客园 - 司徒正美
雷峰网
雷峰网
宝玉的分享
宝玉的分享
钛媒体:引领未来商业与生活新知
钛媒体:引领未来商业与生活新知
博客园_首页
WordPress大学
WordPress大学
Google DeepMind News
Google DeepMind News
GbyAI
GbyAI
MyScale Blog
MyScale Blog
Apple Machine Learning Research
Apple Machine Learning Research
美团技术团队
I
InfoQ
博客园 - Franky
CTFtime.org: upcoming CTF events
CTFtime.org: upcoming CTF events
博客园 - 叶小钗
阮一峰的网络日志
阮一峰的网络日志
Cyberwarzone
Cyberwarzone
C
CXSECURITY Database RSS Feed - CXSecurity.com
S
Schneier on Security
P
Privacy & Cybersecurity Law Blog
T
Threatpost
Cloudbric
Cloudbric
D
Docker
M
MIT News - Artificial intelligence
Recent Commits to openclaw:main
Recent Commits to openclaw:main
Vercel News
Vercel News
Martin Fowler
Martin Fowler
J
Java Code Geeks
AWS News Blog
AWS News Blog
The Cloudflare Blog
酷 壳 – CoolShell
酷 壳 – CoolShell
L
Lohrmann on Cybersecurity
Hacker News: Ask HN
Hacker News: Ask HN
Last Week in AI
Last Week in AI
S
Security @ Cisco Blogs
Help Net Security
Help Net Security
C
Cisco Blogs
V
V2EX
博客园 - 【当耐特】
I
Intezer
爱范儿
爱范儿
F
Fortinet All Blogs
cs.AI updates on arXiv.org
cs.AI updates on arXiv.org
P
Privacy International News Feed
IT之家
IT之家
L
LINUX DO - 最新话题
B
Blog RSS Feed
K
KPMG report finds enterprise disconnect between AI and its ROI | CIO

The Register - On-Prem: Systems

Qualcomm teases agentic CPUs and smartphones Fujitsu says quantum and AI will replace mainframes in 2035 ZTE & China's NCRC partner for smart interventional medicine Core Scientific accelerates crypto-to-AI pivot Meta to use millions of AWS Graviton cores AI now gobbling up power and management chips for servers Tesla stakes AI dreams on Intel's unfinished AI chip SK Hynix breaks ground on Indiana advanced packaging plant Datacenter boom keeps dirty coal plants alive in the US AMD's Ryzen 9 9950X3D2 Dual Edition tested World's blandest man steps down from CEO job to spend more time in tastefully appointed home World's blandest man steps down from CEO job Intel eases reliance on TSMC with 'Merica-made Core Series 3 processors Intel eases reliance on TSMC with Core Series 3 CPUs Guide to GPU virtualization: passthrough, vGPU, and MIG Orbital datacenter startup admits launch economics don't fly AI-powered mainframe exits are a bubble set to pop Cloud-smart strategy helps Interactive meet GenAI demands Oracle taps Bloom for fuel cells to support datacenter binge Oracle taps Bloom for fuel cells to support datacenter binge UK signs Rolls-Royce SMR design deal Japan going back to the future by reviving its chip industry Japan going back to the future by reviving its chip industry AWS ponders selling its home-grown chips by the rack-load Supply chain challenges risk delaying Nvidia's Rubin GPUs Supply chain challenges risk delaying Nvidia's Rubin GPUs Supermicro investigating alleged China chip smuggling Intel trapped in Elon's reality distortion field UALink delivers 2.0 spec before v. 1.0 silicon ships OpenInfra General Manager on sovereignty and kill switches Anthropic reveals $30bn run rate, plan to use new Google TPU Nvidia embraces optical scale-up as copper reaches limits Nvidia embraces optical scale-up as copper reaches limits IBM wants Arm software on its mainframes for AI support AI datacenters create heat islands around them, paper finds Arm says AI agents need a new CPU. Intel doesn't buy it Memory-makers' shares are down. Don't blame Google Memory-makers' shares are down. Don't blame Google US PC shipments to fall 13% as memory and storage crunch hits budget systems US PC shipments to fall 13% as memory costs surge ZTE showcases intelligent computing at CloudFest 2026 Rebellions eyes global expansion with rack-scale AI platform Rebellions eyes global expansion with rack-scale AI platform Enterprise infrastructure is entering an economic reset AMD doubles up on V-Cache with 9950X3D2 Dual Edition Apple's making more iPhone bits in US, but not the iPhone Three more charged with trying to smuggle GPUs to China Three more charged with trying to smuggle GPUs to China Dell slims Pro laptops, boosts battery and cooling Alibaba delivers RISC-V server chip optimized for Chinese AI Alibaba delivers RISC-V server chip optimized for Chinese AI AI-pilled Arm CEO teases mystery products for $1T TAM Arm rolls its own 136-core AGI CPU to chase AI hype train SoftBank builds AI mega-datacenter on nuke site SoftBank builds AI mega-datacenter on nuke site Chip tester shrugged off ransomware – then came the leak Explainer: AI-ready servers Elon Musk proposes 'Terafab' to level up chip production Australia to datacenter operators: BYO energy or stay home Australia to datacenter operators: BYO energy or stay home Supermicro co-founder charged over $2.5B GPU sales to China Blue Origin applies to launch 51,000 datacenter satellites Blue Origin applies to launch 51,000 datacenter satellites Alibaba has made 470,000 AI chips, admits they’re inferior Decoding Nvidia's Groq-powered LPX and the rest of its new rack systems Your next car might need 300 GB of RAM, and so will autonomous robots Your next car might need 300 GB of RAM, and so will robots It's not a binary choice: Boffin builds ternary CPU Nvidia H200 back on in China, production ramping: Huang Nvidia slaps $20B Groq tech into massive new LPX racks to speed AI response time AI Burning Man happens next week – what to expect at Nvidia GTC 2026 Meta reveals four Broadcom-built custom AI chips, claims some outperform commercial silicon Meta reveals custom AI chips it says beat Nvidia ZTE and Orange Morocco launch Livebox 7 for smart homes Ayar Labs taps Wiwynn to cram 1,024 GPUs into a photonic rack system Ayar Labs, Wiwynn to cram 1,024 GPUs into photonic system ZTE and Whale Cloud Showcase Digital Transformation at MWC AI datacenters may gulp NYC's daily water supply at peak Mystery outage behind JetBlue's request for grounding HPE tweaks T&Cs so it can change quotes as RAM prices rise Supermicro launches probe after staff charged with China export violations
Guide to GPU virtualization: passthrough, vGPU, and MIG
VergeIO VergeIO · 2026-04-16 · via The Register - On-Prem: Systems

PARTNER CONTENT GPU workloads are no longer the exclusive territory of research labs and hyperscalers. Engineering teams, data science groups, healthcare organizations, and financial services businesses are all deploying GPU-accelerated infrastructure for AI inference, simulation, visualization, and virtual desktops. For many IT teams, this is new ground. The hardware is familiar, because NVIDIA GPUs fit in standard server slots. But the software isn't.

GPU virtualization has three distinct models. Each makes different tradeoffs between performance, sharing efficiency, and isolation. Understanding which model fits which workload is the first step. Understanding how to operate them is where most deployments run into trouble.

PCIe passthrough: one GPU, one VM

PCIe passthrough is the simplest GPU virtualization model to understand and the hardest to scale. The hypervisor assigns an entire physical GPU to a single virtual machine. The VM communicates directly with the hardware with no abstraction layer and no sharing. From the VM's perspective, it owns a physical GPU.

Passthrough delivers maximum performance. It is the right choice when a single workload must have the full card: Large model training runs, high-fidelity physics simulations, or rendering pipelines that saturate GPU memory. Applications that require bare-metal GPU behavior and cannot tolerate any virtualization overhead run cleanly in the passthrough model.

The tradeoffs are significant. One physical GPU per VM means utilization collapses the moment that workload finishes. Most platforms don't support live migration of a passthrough VM. Instead, you must stop it first. At scale, passthrough turns expensive GPU hardware into a rigid single-tenant resource with no flexibility for sharing or density.

NVIDIA vGPU: one GPU, multiple VMs

NVIDIA vGPU is a full GPU virtualization software stack. It divides a single physical GPU into virtual GPU instances; each VM gets its own vGPU with dedicated memory and a full NVIDIA driver running inside the guest OS. From the application's perspective, the vGPU looks and behaves like a discrete GPU. Software that requires a certified NVIDIA driver runs without modification.

vGPU is the right model for workloads that need GPU access but don't need an entire card. Virtual desktop infrastructure (VDI) environments where knowledge workers need graphics-capable virtual desktops, development environments where multiple engineers share a server, and inference endpoints where multiple services share GPU memory all fit the vGPU model. It delivers density without sacrificing driver compatibility.

vGPU requires a software license from NVIDIA in addition to the hardware cost, and those licenses are recurring. Most platforms define vGPU profiles (the memory and compute allocation for each VM) at creation time. Changing them requires rebuilding the VM.

MIG: hardware-enforced GPU partitioning

Multi-Instance GPU (MIG) is a hardware capability introduced with NVIDIA's Ampere architecture and extended in Hopper and Blackwell. Where vGPU shares GPU resources in software, MIG partitions the physical GPU in silicon. Each MIG instance gets its own dedicated compute engine, memory controller, and memory bandwidth. Hardware enforces the isolation rather than the driver or the hypervisor.

MIG is the right model when workloads need predictable, guaranteed performance and true fault isolation. If one MIG instance encounters an error or runs a noisy workload, it cannot affect neighboring instances. This distinction matters in multi-tenant environments and in regulated industries where workload isolation is a compliance requirement, not just a preference. Private AI inference, where sensitive data must stay inside organizational boundaries and multiple models run simultaneously, is a primary MIG use case.

MIG instances come in fixed sizes defined by the GPU architecture. You cannot create arbitrary partition geometries. MIG is available on datacenter-class GPUs: NVIDIA's A100, H100, and the RTX PRO 6000 Blackwell Server Edition, among others.

The implementation problem: everything requires the command line

Understanding the three models is straightforward, but deploying and operating them isn't.

GPU virtualization on most hypervisors requires command-line expertise that extends well beyond the typical sysadmin skill set. PCIe passthrough requires IOMMU configuration and device binding at the host level. Mistakes at this stage can cause VM instability or host crashes. vGPU requires the NVIDIA vGPU Manager software installed and running on the hypervisor host, with driver versions precisely matched across the hypervisor, the vGPU software stack, and each guest OS. Version mismatches across those three layers are a leading cause of GPU support tickets. A guest OS update that pulls in a new NVIDIA driver can break vGPU functionality across an entire host.

MIG configuration is where CLI complexity peaks. Configuring MIG on an NVIDIA datacenter GPU requires direct use of nvidia-smi: Selecting a MIG profile, enabling MIG mode, creating GPU instances, creating compute instances within each GPU instance, and then assigning those instances to VMs. Reconfiguring MIG when a workload's resource needs change requires destroying the existing instances and recreating them. Every step happens at the command line, on the host. Getting it wrong typically requires rebooting the GPU or the host.

The consequence is predictable: Organizations that could benefit from MIG's isolation and density advantages avoid it entirely because they don't have a GPU specialist on staff. The capability exists in the hardware, but the operational path to using it stops most teams cold.

Ongoing management compounds the initial deployment challenge. Workload requirements change, more VMs need GPU access, inference models grow in size, and new teams need isolated development environments. On most platforms, reconfiguring GPU allocations to accommodate all this means returning to the command line. There is no unified view of GPU utilization alongside compute and memory in the same dashboard. GPU management is a separate discipline from infrastructure management.

Use cases by GPU virtualization type

PCIe passthrough suits workloads that need the entire GPU and don't need to share it, such as LLM training runs, fluid dynamics simulation, computational chemistry, and high-resolution rendering. These jobs typically run to completion and release the GPU. The low utilization window between jobs is the cost of the passthrough model.

NVIDIA vGPU suits density-focused deployments. These include VDI environments where engineering, design, and scientific visualization teams need GPU-accelerated desktops from centralized infrastructure, AI development environments where multiple developers need GPU access simultaneously, and inference endpoints where multiple models share a high-memory GPU. NVIDIA RTX Virtual Workstation (vWS) supports professional visualization workloads, Virtual PC (vPC) supports knowledge worker desktops, and Virtual Applications (vApps) delivers individual GPU-accelerated applications without full virtual desktop overhead.

MIG suits multi-tenant scenarios where isolation is non-negotiable: Private AI inference where each model runs in a guaranteed, isolated environment; regulated industries where healthcare imaging, financial modeling, or defense workloads require hardware-level separation; and research organizations where multiple teams share expensive GPU infrastructure without contention.

How VergeOS changes the operational equation

VergeOS supports all three GPU virtualization models (PCIe passthrough, NVIDIA vGPU, and MIG) as native capabilities within a single platform. NVIDIA introduced VergeOS as a validated vGPU platform, with joint support across RTX Virtual Workstation (vWS), Virtual PC (vPC), and Virtual Applications (vApps). Validated hardware includes the A100, A30, A40, and L40 series, plus the RTX PRO 6000 Blackwell Server Edition for MIG vGPU.

The operational difference is where VergeOS diverges from the typical deployment story. GPU configuration including MIG is point-and-click in the same interface used for compute, storage, and networking. No nvidia-smi. No command-line steps. Selecting a MIG profile, assigning GPU resources to a VM, and even reconfiguring when workload requirements change all happen through the same interface an IT generalist already operates.

Driver management follows the same principle. Upload a driver once. VergeOS builds the ISO and automatically deploys it to every GPU-enabled VM at assignment. VergeOS replaces the three-layer version management problem covering hypervisor, vGPU software stack, and guest OS with a single upload and automatic distribution.

GPU utilization appears alongside CPU and memory in the same monitoring dashboard. There is no separate GPU management plane, no additional tool to learn, and no specialist required to stand up and operate the platform.

Organizations with existing VergeOS deployments add GPU capabilities by installing supported NVIDIA hardware in their cluster nodes. VergeOS detects the hardware automatically. The same platform that manages your storage and networking manages your GPU infrastructure without a learning curve that most IT teams can't afford.

Watch the GPU Virtualization Without the Complexity on-demand webinar for a live demonstration of all three GPU modes in the VergeOS interface. Download the GPU Virtualization Without the Complexity white paper for a full technical breakdown of GPU modes, driver management, and deployment scenarios.

Contributed by VergeIO.