





















Abstract:Gradient-based optimizers are highly sensitive to design choices in their adaptive learning rate mechanisms. To address this limitation, we introduce POP, a meta-learned Reinforcement Learning (RL) policy that predicts adaptive learning rates for gradient descent, conditioned on the contextual information provided in the optimization trajectory. Our method introduces a novel RL reward formulation, a new function-scaling strategy for in-distribution generalization, and a novel prior that is used to sample millions of synthetic optimization problems. We evaluate POP on an established benchmark including 43 optimization functions of various complexity, where it significantly outperforms gradient-based methods. Our evaluation demonstrates strong generalization capabilities without task-specific tuning.
| Comments: | Under Review |
| Subjects: | Machine Learning (cs.LG) |
| Cite as: | arXiv:2602.15473 [cs.LG] |
| (or arXiv:2602.15473v2 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2602.15473 arXiv-issued DOI via DataCite |
From: Jan Kobiolka [view email]
[v1]
Tue, 17 Feb 2026 10:27:07 UTC (4,728 KB)
[v2]
Tue, 12 May 2026 12:43:04 UTC (10,427 KB)
此内容由惯性聚合(RSS阅读器)自动聚合整理,仅供阅读参考。 原文来自 — 版权归原作者所有。