





























Abstract:Continuous knowledge updating for pre-trained large language models (LLMs) is increasingly necessary yet remains challenging. Although inference-time methods like In-Context Learning (ICL) and Retrieval-Augmented Generation (RAG) are popular, they face constraints in context budgets, costs, and retrieval fragmentation. Departing from these context-dependent paradigms, this work investigates a parametric approach using Low-Rank Adaptation (LoRA) as a modular knowledge memory. Although few recent works examine this concept, the fundamental mechanics governing its capacity and composability remain largely unexplored. We bridge this gap through the first systematic empirical study mapping the design space of LoRA-based memory, ranging from characterizing storage capacity and optimizing internalization to scaling multi-module systems and evaluating long-context reasoning. Rather than proposing a single architecture, we provide practical guidance on the operational boundaries of LoRA memory. Overall, our findings position LoRA as the complementary axis of memory alongside RAG and ICL, offering distinct advantages.
| Comments: | ICML 2026 |
| Subjects: | Machine Learning (cs.LG) |
| Cite as: | arXiv:2603.01097 [cs.LG] |
| (or arXiv:2603.01097v2 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2603.01097 arXiv-issued DOI via DataCite |
|
| Journal reference: | Proceedings of the Forty-Third International Conference on Machine Learning (ICML), 2026 |
From: Seungju Back [view email]
[v1]
Sun, 1 Mar 2026 13:28:57 UTC (563 KB)
[v2]
Wed, 6 May 2026 01:33:20 UTC (563 KB)
此内容由惯性聚合(RSS阅读器)自动聚合整理,仅供阅读参考。 原文来自 — 版权归原作者所有。