





















Abstract:We introduce State Vector Space Partitioning (SVSP), a novel method to mimic a black box reinforcement learning policy using a set of human-interpretable subpolicies. By partitioning a distillation dataset of state action pairs with linear support vector machine splits, SVSP constructs a compact and structured representation of the original policy. Our method improves mean return by +7.4\% over previous critic driven state partitioning attempts such as Voronoi State Partitioning (VSP) and +2.8\% over the original TD3 policy, while reducing the number of required subpolicies against VSP by 82.1\%. Our results pave the path towards a more flexible form of distillation where both the decision boundary and surrogate models can be chosen within a margin of the original black box behavior.
| Comments: | Accepted for poster presentation at HHAI 2026 |
| Subjects: | Machine Learning (cs.LG); Human-Computer Interaction (cs.HC) |
| Cite as: | arXiv:2605.04254 [cs.LG] |
| (or arXiv:2605.04254v1 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2605.04254 arXiv-issued DOI via DataCite (pending registration) |
From: Senne Deproost [view email]
[v1]
Tue, 5 May 2026 19:40:05 UTC (76 KB)
此内容由惯性聚合(RSS阅读器)自动聚合整理,仅供阅读参考。 原文来自 — 版权归原作者所有。