


















Abstract:In the modern age of large-scale AI, federated learning has become an increasingly important tool for training large populations of AI agents; however, its computational and communication costs can rapidly fail to scale with the number of agents. This is precisely where decentralized agentic strategies shine: each agent acts autonomously, using only its own state together with a minimal summary of the ensemble, namely the mean-field. We derive the unique optimal decentralized policy in closed form. Optimality is characterized through a worst-client/minimax criterion: minimizing the under-performer regret, namely the maximal online cost incurred by the weakest agent in the ensemble. We further prove that the resulting decentralized policy asymptotically converges, in the large-population limit, to the Nash-optimal centralized policy, whose direct computation is not scalable. We use an online weighting mechanism to optimize the server-computed mixture of client predictions, thereby improving the mean prediction in addition to the previously optimized weakest-client prediction. Numerical experiments verify our theoretical guarantees and demonstrate that our decentralized policy typically outperforms natural greedy decentralized baselines.
| Comments: | 43 pages, 11 tables, 1 figure |
| Subjects: | Machine Learning (cs.LG) |
| MSC classes: | 91A80, 91A16, 93E20, 49N10 |
| ACM classes: | I.2.11; I.2.8 |
| Cite as: | arXiv:2605.05492 [cs.LG] |
| (or arXiv:2605.05492v1 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2605.05492 arXiv-issued DOI via DataCite (pending registration) |
From: Xuwei Yang [view email]
[v1]
Wed, 6 May 2026 22:26:59 UTC (200 KB)
此内容由惯性聚合(RSS阅读器)自动聚合整理,仅供阅读参考。 原文来自 — 版权归原作者所有。