












Abstract:Directional-statistics-based mask-based blind speech separation (BSS) clusters normalized time-frequency (TF) observations from $M$ microphones on the complex unit sphere, without relying on plane-wave or spherical-wave assumptions. Existing methods use separately defined angular mixture models, which makes the effect of density shape difficult to isolate. This paper proposes a complex spherical Student's $t$ mixture model (cSTMM) that connects the complex angular central Gaussian mixture model (cACGMM), complex Bingham mixture model (cBMM), and complex Watson mixture model (cWMM) through the degrees of freedom $\nu$ and eigenvalue constraints. We derive a latent-scale expectation-maximization (EM) framework with an approximate M-step based on high-concentration approximation (HCA). On noise-free LibriSpeech mixtures reverberated using measured room impulse responses (RIRs), the development-selected value $\nu^\ast=1$ outperformed the cACGMM-equivalent choice $\nu=M$ in all 18 test conditions, yielding a mean signal-to-distortion ratio improvement (SDRi) gain of $0.25\,\mathrm{dB}$. The model reduces to the cACGMM at $\nu=M$ and approaches the cBMM/cWMM in the large-$\nu$ limits.
From: Nobutaka Ito B.E. M.E. Ph.D. [view email]
[v1]
Mon, 25 May 2026 07:15:39 UTC (328 KB)
[v2]
Thu, 6 Aug 2026 07:16:51 UTC (48 KB)
此内容由惯性聚合(RSS阅读器)自动聚合整理,仅供阅读参考。 原文来自 — 版权归原作者所有。