



















Abstract:We study stochastic minimum-cost reach-avoid reinforcement learning, where an agent must satisfy a reach-avoid specification with probability at least $p$ while minimizing expected cumulative costs in stochastic environments. Existing safe and constrained reinforcement learning methods typically fail to jointly enforce probabilistic reach-avoid constraints and optimize cost in the learning setting in stochastic environments. To address this challenge, we introduce reach-avoid probability certificates (RAPCs), which identify states from which stochastic reach-avoid constraints are satisfiable. Building on RAPCs, we develop a contraction-based Bellman formulation that serves as a principled surrogate for integrating reach-avoid considerations into reinforcement learning, enabling cost optimization under probabilistic constraints. We establish almost sure convergence of the proposed algorithms to locally optimal policies with respect to the resulting objective. Experiments in the MuJoCo simulator demonstrate improved cost performance and consistently higher reach-avoid satisfaction rates.
| Comments: | Accepted at the Forty-third International Conference on Machine Learning (ICML 2026) |
| Subjects: | Machine Learning (cs.LG) |
| Cite as: | arXiv:2605.11975 [cs.LG] |
| (or arXiv:2605.11975v1 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2605.11975 arXiv-issued DOI via DataCite (pending registration) |
From: Jingduo Pan [view email]
[v1]
Tue, 12 May 2026 11:31:36 UTC (2,139 KB)
此内容由惯性聚合(RSS阅读器)自动聚合整理,仅供阅读参考。 原文来自 — 版权归原作者所有。