






















Abstract:The deployment of intelligent reinforcement learning (RL) agents on resource-constrained edge devices remains a fundamental challenge due to the substantial memory, computational, and energy requirements of modern deep learning systems. While large language models (LLMs) have emerged as powerful architectures for decision-making agents, their multi-billion parameter scale confines them to cloud-based deployment, raising concerns about latency, privacy, and connectivity dependence.
We introduce BitRL, a framework for building RL agents using 1-bit quantized language models that enables practical on-device learning and inference under severe resource constraints. Leveraging the BitNet b1.58 architecture with ternary weights (-1, 0, +1) and an optimized inference stack, BitRL achieves 10-16x memory reduction and 3-5x energy efficiency improvements over full-precision baselines while maintaining 85-98 percent of task performance across benchmarks.
We provide theoretical analysis of quantization as structured parameter perturbation, derive convergence bounds for quantized policy gradients under frozen-backbone architectures, and identify the exploration-stability trade-off in extreme quantization. Our framework systematically integrates 1-bit quantized language models with reinforcement learning for edge deployment and demonstrates effectiveness on commodity hardware.
| Comments: | 6pages, 1 Figure, IEEE International Conference of Frontiers of Engineering and Emerging Technologies 2026 |
| Subjects: | Machine Learning (cs.LG) |
| Cite as: | arXiv:2604.24273 [cs.LG] |
| (or arXiv:2604.24273v1 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2604.24273 arXiv-issued DOI via DataCite (pending registration) |
From: Mohammad Sakib Mahmood [view email]
[v1]
Mon, 27 Apr 2026 10:03:37 UTC (423 KB)
此内容由惯性聚合(RSS阅读器)自动聚合整理,仅供阅读参考。 原文来自 — 版权归原作者所有。