



















Abstract:A central challenge in multi-task reinforcement learning (RL) is to train generalist policies capable of performing tasks not seen during training. To facilitate such generalization, linear temporal logic (LTL) has emerged as a powerful formalism for specifying structured, temporally extended tasks to RL agents. While existing approaches to LTL-guided multi-task RL demonstrate generalization across LTL specifications, they are unable to generalize to unseen vocabularies of propositions (or "symbols"), which describe high-level events in LTL. We present PlatoLTL, a novel approach that enables policies to zero-shot generalize not only compositionally across LTL structures, but also parametrically across propositions. We model propositions as parameterized instances of atomic predicates, allowing policies to learn shared structure across related propositions. We propose a novel architecture that embeds and composes parameterized propositions to represent LTL formulae, and demonstrate zero-shot generalization in a range of challenging environments.
| Comments: | 14 pages, 4 figures (main paper). 22 pages, 11 figures (appendix) |
| Subjects: | Machine Learning (cs.LG) |
| Cite as: | arXiv:2601.22891 [cs.LG] |
| (or arXiv:2601.22891v2 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2601.22891 arXiv-issued DOI via DataCite |
From: Jacques Cloete [view email]
[v1]
Fri, 30 Jan 2026 12:11:55 UTC (3,372 KB)
[v2]
Thu, 7 May 2026 12:55:00 UTC (3,523 KB)
此内容由惯性聚合(RSS阅读器)自动聚合整理,仅供阅读参考。 原文来自 — 版权归原作者所有。