












Abstract:This paper investigates hybrid real- and complex-valued neural networks for monaural speech enhancement. While complex-valued models can process time-frequency representations natively, they often increase computational cost and can be inefficient in small-model regimes. We therefore study a matched-parameter hybrid architecture that combines a real-valued magnitude-mask branch with a complex-valued additive correction branch, coupled via domain conversion functions at the bottleneck. The approach is applied to convolutional denoising autoencoder and convolutional-recurrent network architectures. Averaged over four SNRs, the hybrid models improve intelligibility and quality over the considered counterparts, while substantially reducing the number of operations.
From: Luan Vinícius Fiorio [view email]
[v1]
Thu, 25 Sep 2025 14:00:57 UTC (138 KB)
[v2]
Wed, 12 Aug 2026 13:43:37 UTC (199 KB)
此内容由惯性聚合(RSS阅读器)自动聚合整理,仅供阅读参考。 原文来自 — 版权归原作者所有。