


























Abstract:We study reward poisoning attacks in reinforcement learning (RL), where an adversary manipulates rewards within constrained budgets to force the target RL agent to adopt a policy that aligns with the attacker's objectives. Prior works on reward poisoning mainly focused on sufficient conditions to design a successful attacker, while only a few studies discussed the infeasibility of targeted attacks. This paper provides the first precise necessity and sufficiency characterization of the attackability of a linear MDP under reward poisoning attacks. Our characterization draws a bright line between the vulnerable RL instances, and the intrinsically robust ones which cannot be attacked without large costs even running vanilla non-robust RL algorithms. Our theory extends beyond linear MDPs -- by approximating deep RL environments as linear MDPs, we show that our theoretical framework effectively distinguishes the attackability and efficiently attacks the vulnerable ones, demonstrating both the theoretical and practical significance of our characterization.
| Subjects: | Machine Learning (cs.LG) |
| Cite as: | arXiv:2604.10062 [cs.LG] |
| (or arXiv:2604.10062v2 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2604.10062 arXiv-issued DOI via DataCite |
From: Haoyang Hong [view email]
[v1]
Sat, 11 Apr 2026 06:55:45 UTC (20,406 KB)
[v2]
Tue, 14 Apr 2026 22:53:06 UTC (20,406 KB)
此内容由惯性聚合(RSS阅读器)自动聚合整理,仅供阅读参考。 原文来自 — 版权归原作者所有。