












Abstract:Speech enhancement (SE) models advance rapidly, yet how input degradation affects their internal representations remains underexplored. We introduce a probing framework to characterize how internal representations in SE models behave under controlled input degradation. We probe three SE models across controlled levels of signal-to-noise ratio (SNR) and reverberation, quantified by $C_{50}$, measuring layer-wise similarity to clean references with Centered Kernel Alignment (CKA) and summarizing each layer by a linear fit against degradation level: the intercept measures robustness, whereas the slope measures sensitivity. All three models are sharply non-uniform across depth, but they organize that non-uniformity differently. MUSE and MP-SENet grow more sensitive with depth, the sharpest transitions falling at MUSE's skip-connection junctions, where encoder information is reintegrated; Demucs inverts the trend. A randomly initialized model shows a near-flat profile, with slopes one to two orders of magnitude smaller, and the profile forms during fine-tuning, indicating that it is induced by the enhancement objective rather than a particular model design. Because CKA saturates at the clean reference, intercept and slope are partly coupled; we derive the identity relating them and report a \emph{saturation spread} statistic that indicates when their relationship is informative. Together, these results characterize where SE models are most sensitive to degradation. An exploratory analysis of whether residual variation tracks output-level quality, after controlling for SNR, shows that the speaker, rather than the utterance, must be treated as the sampling unit. Code and precomputed analysis artifacts for the main sweeps are publicly available.
From: Yair Amar [view email]
[v1]
Sat, 29 Nov 2025 13:29:00 UTC (149 KB)
[v2]
Sun, 3 May 2026 20:26:20 UTC (1,041 KB)
[v3]
Mon, 14 Sep 2026 18:08:48 UTC (309 KB)
此内容由惯性聚合(RSS阅读器)自动聚合整理,仅供阅读参考。 原文来自 — 版权归原作者所有。