
























Abstract:Learning in environments with sparse rewards remains a fundamental challenge in reinforcement learning. Artificial curiosity addresses this limitation through intrinsic rewards to guide exploration, however, the precise formulation of these rewards has remained elusive. Ideally, such rewards should depend on the agent's information about the environment, remaining agnostic to its representation -- an invariance central to information geometry. Leveraging this, we show that information monotonicity and invariance under the agent-environment interaction uniquely constrains intrinsic rewards to strictly concave functions of the reciprocal occupancy. Requiring these rewards to yield a principled exploration-exploitation trade-off, via information geodesic interpolation on the occupancy manifold, effectively limits the candidates to those determined by a scalar parameter. Remarkably, special values of this parameter are found to correspond to count-based and maximum entropy exploration. This framework provides important constraints to the engineering of intrinsic rewards while integrating foundational exploration methods into a single, cohesive model.
| Comments: | Comments: 24 pages, 2 figures; version accepted for publication at AISTATS 2026 |
| Subjects: | Machine Learning (cs.LG) |
| Cite as: | arXiv:2504.06355 [cs.LG] |
| (or arXiv:2504.06355v2 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2504.06355 arXiv-issued DOI via DataCite |
From: Alexander Nedergaard [view email]
[v1]
Tue, 8 Apr 2025 18:04:15 UTC (172 KB)
[v2]
Fri, 17 Apr 2026 09:05:35 UTC (298 KB)
此内容由惯性聚合(RSS阅读器)自动聚合整理,仅供阅读参考。 原文来自 — 版权归原作者所有。