






























Abstract:A key capability of intelligent agents is operating under partial observability: reasoning and acting effectively despite missing or incomplete state observations. While recurrent (memory-based) policies learned via reinforcement learning address this by encoding history into latent state representations, their internal dynamics remain uninterpretable black boxes. This paper establishes a formal link between these hidden states and the Pontryagin minimum principle (PMP) from optimal control. We demonstrate that for standard recurrent architectures, latent representations map directly to PMP co-states, which allows the readout layer to be interpreted as performing Hamiltonian minimization. Because standard reward maximization does not naturally discover this alignment, we introduce a PMP-derived co-state loss to explicitly structure the internal dynamics. Empirically, this approach matches or improves performance on partially observable DMControl tasks, and is robust against zero-shot out-of-distribution sensor masking. By framing recurrent networks as dynamic processes governed by the minimum principle, we provide a principled approach to designing robust continuous control policies.
| Comments: | 17 pages, 5 figures |
| Subjects: | Machine Learning (cs.LG) |
| Cite as: | arXiv:2605.05373 [cs.LG] |
| (or arXiv:2605.05373v1 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2605.05373 arXiv-issued DOI via DataCite |
From: David Leeftink [view email]
[v1]
Wed, 6 May 2026 18:53:33 UTC (2,870 KB)
此内容由惯性聚合(RSS阅读器)自动聚合整理,仅供阅读参考。 原文来自 — 版权归原作者所有。