






















Abstract:While Multi-Agent Reinforcement Learning (MARL) algorithms achieve unprecedented successes across complex continuous domains, their standard deployment strictly adheres to a synchronous operational paradigm. Under this paradigm, agents are universally forced to execute deep neural network inferences at every micro-frame, regardless of immediate necessity. This dense throughput acts as a fundamental barrier to physical deployment on edge-devices where thermal and metabolic budgets are highly constrained. We propose Epistemic Time-Dilation MAPPO (ETD-MAPPO), augmented with a Dual-Gated Epistemic Trigger. Instead of depending on rigid frame-skipping (macro-actions), agents autonomously modulate their execution frequency by interpreting aleatoric uncertainty (via Shannon entropy of their policy) and epistemic uncertainty (via state-value divergence in a Twin-Critic architecture). To format this, we structure the environment as a Semi-Markov Decision Process (SMDP) and build the SMDP-Aligned Asynchronous Gradient Masking Critic to ensure proper credit assignment. Empirical findings demonstrate massive improvements (> 60% relative baseline acquisition leaps) over current temporal models. By assessing LBF, MPE, and the 115-dimensional state space of Google Research Football (GRF), ETD correctly prevented premature policy collapse. Remarkably, this unconstrained approach leads to emergent Temporal Role Specialization, reducing computational overhead by a statistically dominant 73.6% entirely during off-ball execution without deteriorating centralized task dominance.
| Comments: | 14 pages, 5 figures. Code available at: this https URL. Related materials available on Zenodo: https://doi.org/10.5281/zenodo.19206838 |
| Subjects: | Multiagent Systems (cs.MA); Machine Learning (cs.LG) |
| ACM classes: | I.2.11 |
| Cite as: | arXiv:2603.23722 [cs.MA] |
| (or arXiv:2603.23722v2 [cs.MA] for this version) | |
| https://doi.org/10.48550/arXiv.2603.23722 arXiv-issued DOI via DataCite |
From: Igor Jankowski [view email]
[v1]
Tue, 24 Mar 2026 21:19:06 UTC (1,501 KB)
[v2]
Tue, 19 May 2026 05:50:46 UTC (1,500 KB)
此内容由惯性聚合(RSS阅读器)自动聚合整理,仅供阅读参考。 原文来自 — 版权归原作者所有。