





















Abstract:Vanishing gradients and overfitting are central problems in machine learning, yet are typically analyzed in asymptotic regimes that obscure their dynamical origins. Here we provide a dynamical description of learning in multi-layer perceptrons (MLPs) via a minimal model inspired by Fukumizu and Amari. We show that training dynamics traverse plateau and near-optimal regions, both organized by saddle structures, before converging to an overfitting regime. Under suitable conditions on the data, this regime collapses to a single attractor modulo symmetry. Furthermore, for finite noisy datasets, convergence to the theoretical optimum is impossible, and the dynamics necessarily settle into an overfitting solution.
| Subjects: | Machine Learning (cs.LG); Adaptation and Self-Organizing Systems (nlin.AO) |
| Cite as: | arXiv:2604.02393 [cs.LG] |
| (or arXiv:2604.02393v2 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2604.02393 arXiv-issued DOI via DataCite |
From: Alex Maleknia [view email]
[v1]
Thu, 2 Apr 2026 11:40:53 UTC (2,259 KB)
[v2]
Fri, 17 Apr 2026 08:31:38 UTC (2,259 KB)
此内容由惯性聚合(RSS阅读器)自动聚合整理,仅供阅读参考。 原文来自 — 版权归原作者所有。