













Abstract:Modern multi-tenant AI clusters are increasingly communication-bound, driven by high-volume and multi-round GPU-to-GPU collective communication. Consequently, the GPU dispatcher's choice of a physical GPU subset for each tenant largely determines the job's effective collective bandwidth and thus its performance ceiling. Existing dispatchers predominantly rely on static, topology-aware heuristics that prioritize GPU resource compactness, assuming that minimizing physical distance maximizes communication bandwidth.
However, we reveal that this assumption often fails due to complex system-level bottlenecks, such as non-linear NIC saturation and inter-node link heterogeneity. This paper presents BandPilot, a performance- and contention-aware GPU dispatching primitive that optimizes effective collective bandwidth for multi-tenant AI clusters. Specifically, BandPilot learns a data-efficient bandwidth model from sparse NCCL measurements via a hierarchical design. Guided by the model, BandPilot uses an equilibrium-driven heuristic as a fast front end, and invokes a pruned elimination search when a controller predicts that further refinement is worthwhile. To account for multi-tenant interference, BandPilot virtually merges a candidate allocation with co-located cross-host jobs to conservatively estimate shared bottleneck capacity and predict contention-degraded bandwidth. Across a 32-GPU H100 cluster and heterogeneous simulations, BandPilot achieves 90-97% bandwidth efficiency relative to the best-found reference, improving average efficiency by 20-30% over topology-compactness heuristics.
From: Kunming Zhang [view email]
[v1]
Wed, 18 Jun 2025 16:10:17 UTC (782 KB)
[v2]
Wed, 25 Jun 2025 08:27:45 UTC (977 KB)
[v3]
Fri, 22 Aug 2025 02:33:47 UTC (1,635 KB)
[v4]
Tue, 6 Jan 2026 12:13:40 UTC (3,074 KB)
[v5]
Sun, 16 Aug 2026 15:01:17 UTC (4,406 KB)
此内容由惯性聚合(RSS阅读器)自动聚合整理,仅供阅读参考。 原文来自 — 版权归原作者所有。