










Abstract:Fluid antenna systems (FASs) improve wireless links by repositioning antenna elements to exploit favorable spatial channel variations. Jointly optimizing fluid antenna (FA) positions, beamforming, and transmit power in a multi-cell network is challenging because the problem is non-convex and each base station has only local information during decentralized execution. However, the representative multi-agent reinforcement learning (MARL) algorithms, namely the on-policy multi-agent proximal policy optimization (MAPPO) and the off-policy multi-agent twin delayed deep deterministic policy gradient (MATD3), suffer from excessively long training times. To address this challenge, we formulate the problem as a decentralized partially observable Markov decision process (Dec-POMDP) and propose multi-agent group relative policy optimization (MAGRPO) under centralized training with decentralized execution. MAGRPO constructs relative advantages from groups of joint trajectories, thereby eliminating the centralized critic and generalized advantage estimation used by MAPPO; under parameter sharing, this critic-free design reduces the per-step computational complexity by approximately half. Simulations show that joint FA optimization provides several-fold sum-rate gains over fixed-position configurations. Across the evaluated antenna settings, MAGRPO achieves test sum rates higher than those of MATD3 and comparable to or slightly higher than those of MAPPO. Meanwhile, it reduces the training time by 20%-23% compared with MAPPO and by about 35% compared with MATD3.
From: Tong Zhang [view email]
[v1]
Sun, 19 Apr 2026 11:05:51 UTC (5,752 KB)
[v2]
Fri, 1 May 2026 12:51:09 UTC (5,753 KB)
[v3]
Thu, 17 Sep 2026 10:03:35 UTC (4,989 KB)
此内容由惯性聚合(RSS阅读器)自动聚合整理,仅供阅读参考。 原文来自 — 版权归原作者所有。