




















Abstract:This paper investigates robust representation learning in offline goal-conditioned reinforcement learning (GCRL). Particularly in sparse reward scenarios, learning representations that align state and goal latents is a challenge that frequently culminates in representation divergence where the encoder drifts toward a low-dimensional, goal-agnostic subspace that destabilizes policy learning. We address this issue by showing that an agent must acquire a fundamental understanding of its environment across multiple scales, from local physical dynamics to long-horizon goal-directed structure. Building on this insight, we propose this http URL, a framework that leverages multi-scale predictive supervision to enforce goal-directed alignment within the latent space. We demonstrate that this http URL leads to improved representation quality and strong performance on both vision and state-based tasks. Furthermore, we show that our approach is exceptionally resilient under realistic, challenging data regimes, maintaining state-of-the-art performance across a wide variety of tasks, trajectory stitching scenarios, and extreme noise conditions.
| Subjects: | Machine Learning (cs.LG) |
| Cite as: | arXiv:2605.09364 [cs.LG] |
| (or arXiv:2605.09364v1 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2605.09364 arXiv-issued DOI via DataCite (pending registration) |
From: Valliappan Chidambaram Adaikkappan [view email]
[v1]
Sun, 10 May 2026 06:27:20 UTC (2,737 KB)
此内容由惯性聚合(RSS阅读器)自动聚合整理,仅供阅读参考。 原文来自 — 版权归原作者所有。