























Abstract:Stochastic Gradient Descent (SGD) is a known stochastic iterative method popular for large-scale convex optimization problems due to its simple implementation and scalability. Some objectives, such as those found in complex-valued neural networks, benefit from updates like in SGD and Gradient Descent (GD) with a newly defined ``gradient'' that allows for complex parameters. This complex variant of the SGD/GD methods has already been proposed, but convergence guarantees without analyticity constraints have not yet been provided. We propose a variant of SGD (complex SGD) that allows for complex parameters, and we provide convergence guarantees under assumptions that parallel those from the real setting. Notably, these results extend to GD as well, and with the same set of assumptions, we confirm that some directional bias results extend from the real to the complex setting for kernel regression problems. We provide empirical results demonstrating the efficacy of the complex SGD in kernel regression problems utilizing complex reproducing kernel Hilbert spaces. In particular, we demonstrate we may recover superoscillation functions and Blaschke products from the Fock Space and Hardy Space, respectively, as the optimal functions for a particular choice of a loss function.
| Subjects: | Machine Learning (cs.LG); Complex Variables (math.CV); Numerical Analysis (math.NA) |
| MSC classes: | 65K10, 62L20, 46E22, 30H10, 65F10, 90C25 |
| Cite as: | arXiv:2604.23017 [cs.LG] |
| (or arXiv:2604.23017v1 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2604.23017 arXiv-issued DOI via DataCite (pending registration) |
From: Natanael Alpay [view email]
[v1]
Fri, 24 Apr 2026 21:08:39 UTC (257 KB)
此内容由惯性聚合(RSS阅读器)自动聚合整理,仅供阅读参考。 原文来自 — 版权归原作者所有。