Posterior Sampling Reinforcement Learning with Gaussian Processes for Continuous Control: Sublinear Regret Bounds for Unbounded State Spaces
[Submitted on 9 Mar 2026 (v1), last revised 22 Jul 2026 (this ve
·
2026-03-09
·
via stat.ML updates on arXiv.org
arXiv:2603.08287v3 Announce Type: replace Abstract: We analyze the Bayesian regret of the Gaussian process po…
此内容由惯性聚合(RSS阅读器)自动聚合整理,仅供阅读参考。 原文来自 — 版权归原作者所有。