





















Abstract:Multi-objective reinforcement learning (MORL) allows a user to express preference over outcomes in terms of the relative importance of the objectives, but standard metrics cannot capture whether changes in preference reliably change the agent's behavior in the intended way, a property termed controllability. As a result, preference-conditioned agents can score well on standard MORL metrics while being insensitive to the preference input. If the ability to control agents cannot be reliably assessed, the symbolic interface that MORL provides between user intent and agent behavior is broken. Mainstream MORL metrics alone fail to measure the controllability of preference-conditioned agents, motivating a complementary metric specifically designed to that end. We hope the results spur discussion in the community on existing evaluation protocols to consolidate advances in preference adaptation in MORL to larger and more complex problems.
| Subjects: | Machine Learning (cs.LG) |
| Cite as: | arXiv:2605.10585 [cs.LG] |
| (or arXiv:2605.10585v1 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2605.10585 arXiv-issued DOI via DataCite (pending registration) |
From: Pau De Las Heras Molins [view email]
[v1]
Mon, 11 May 2026 13:58:49 UTC (3,286 KB)
此内容由惯性聚合(RSS阅读器)自动聚合整理,仅供阅读参考。 原文来自 — 版权归原作者所有。