





















Abstract:Offline reinforcement learning (RL) enables learning effective policies from fixed datasets without any environment interaction. Existing methods typically employ policy constraints to mitigate the distribution shift encountered during offline RL training. However, because the scale of the constraints varies across tasks and datasets of differing quality, existing methods must meticulously tune hyperparameters to match each dataset, which is time-consuming and often impractical. We propose Adaptive Scaling of Policy Constraints (ASPC), a second-order differentiable framework that dynamically balances RL and behavior cloning (BC) during training. We theoretically analyze its performance improvement guarantee. In experiments on 39 datasets across four D4RL domains, ASPC using a single hyperparameter configuration outperforms other adaptive constraint methods and state-of-the-art offline RL algorithms that require per-dataset tuning while incurring only minimal computational overhead. The code will be released at this https URL.
| Subjects: | Machine Learning (cs.LG) |
| Cite as: | arXiv:2508.19900 [cs.LG] |
| (or arXiv:2508.19900v2 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2508.19900 arXiv-issued DOI via DataCite |
From: Jing Tan [view email]
[v1]
Wed, 27 Aug 2025 14:00:18 UTC (1,376 KB)
[v2]
Wed, 29 Apr 2026 09:32:34 UTC (1,809 KB)
此内容由惯性聚合(RSS阅读器)自动聚合整理,仅供阅读参考。 原文来自 — 版权归原作者所有。