





















Abstract:Diffusion models have achieved remarkable success, yet their training remains inefficient due to a severe optimization bottleneck, which we term Representation Degradation. As noise levels increase, the outputs of the trained model exhibit progressive structural distortion, which can destabilize training and impair generation quality. Our analysis suggests that this instability is driven by mismatched target recoverability, which is associated with Neural Tangent Kernel (NTK) spectral weakening and effective low-rank behavior. To address this, we propose Elucidated Representation Diffusion (ERD), a plug-and-play framework that dynamically reallocates optimization effort according to effective recoverability. By stabilizing representation learning without external supervision, ERD accelerates convergence and achieves strong empirical performance across diffusion backbones.
| Subjects: | Machine Learning (cs.LG) |
| Cite as: | arXiv:2605.10790 [cs.LG] |
| (or arXiv:2605.10790v1 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2605.10790 arXiv-issued DOI via DataCite (pending registration) |
From: Zhipeng Yao [view email]
[v1]
Mon, 11 May 2026 16:21:45 UTC (28,088 KB)
此内容由惯性聚合(RSS阅读器)自动聚合整理,仅供阅读参考。 原文来自 — 版权归原作者所有。