
























Abstract:Multi-agent coordination under partial observability requires agents to share complementary private information. While recent methods optimize messages for intermediate objectives (e.g., reconstruction accuracy or mutual information), rather than decision quality, we introduce \textbf{SeqComm-DFL}, unifying the sequential communication with decision-focused learning for task performance. Our approach features \emph{value-aware message generation with sequential Stackelberg conditioning}: messages maximize receiver decision quality and are generated in priority order, with agents conditioning on their predecessors. The \emph{guidance potential} determined by their prosocial ordering. We extend Optimal Model Design to communication-augmented world models with QMIX factorization, enabling efficient end-to-end training via implicit differentiation. We prove information-theoretic bounds showing that communication value scales with coordination gaps and establish $\mathcal{O}(1/\sqrt{T})$ convergence for the bilevel optimization, where $T$ denotes the number of training iterations. On collaborative healthcare and StarCraft Multi-Agent Challenge (SMAC) benchmarks, SeqComm-DFL achieves four to six times higher cumulative rewards and over 13\% win rate improvements, enabling coordination strategies inaccessible under information asymmetry.
| Comments: | 15 pages, 6 figures, 3 tables. Includes appendix. Submitted to ICML 2026. Code available at this https URL |
| Subjects: | Machine Learning (cs.LG); Multiagent Systems (cs.MA) |
| MSC classes: | 68T05, 90C15, 68W15 |
| ACM classes: | I.2.6; I.2.11; F.2.2 |
| Cite as: | arXiv:2604.08944 [cs.LG] |
| (or arXiv:2604.08944v1 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2604.08944 arXiv-issued DOI via DataCite |
From: Wesley Marrero [view email]
[v1]
Fri, 10 Apr 2026 04:23:29 UTC (209 KB)
此内容由惯性聚合(RSS阅读器)自动聚合整理,仅供阅读参考。 原文来自 — 版权归原作者所有。