












Abstract:Modeling speech variation is key to natural, expressive generation. Speaker embeddings are commonly used to condition personalized speech systems, but they are typically trained for speaker recognition, where intra-speaker variability is suppressed and inter-speaker separation is maximized. This objective leads to overly compact representations that may discard variations crucial for generation. We revisit this design choice and propose a sub-center modeling framework for speaker embeddings. Instead of a single prototype per speaker, we learn multiple sub-centers during discriminative training, allowing utterances to align with different prototypes. This strategy preserves structured intra-speaker variability while maintaining discriminability. In zero-shot voice conversion, our method improves intelligibility, increases pitch variability, achieves higher naturalness ratings, and retains strong speaker verification performance.
From: Ismail Rasim Ulgen [view email]
[v1]
Fri, 5 Jul 2024 06:54:24 UTC (1,245 KB)
[v2]
Fri, 23 May 2025 20:58:46 UTC (1,958 KB)
[v3]
Thu, 18 Sep 2025 20:22:33 UTC (232 KB)
[v4]
Fri, 28 Aug 2026 02:51:15 UTC (376 KB)
此内容由惯性聚合(RSS阅读器)自动聚合整理,仅供阅读参考。 原文来自 — 版权归原作者所有。