










Abstract:Natural and human-like speech depends on the coordination between prosodic timing and acoustic realization: duration modeling shapes rhythmic structure, while waveform generation determines whether that structure is rendered naturally. In natural speech, duration patterns vary across linguistic contexts and speakers, requiring a TTS system both to capture this variability and to faithfully realize it in the waveform. To address these challenges, we propose FNH-TTS, a VITS-based end-to-end system that jointly improves duration modeling and waveform generation. A mixture-of-experts duration predictor (MoE-DP) uses multiple experts and routing jointly conditioned on linguistic context and speaker information to model diverse duration patterns. For waveform generation, we adopt an inverse short-time Fourier transform (ISTFT)-based generator, providing a more direct and efficient synthesis path. We further employ multi-resolution and sub-band discriminators for fine-grained temporal and spectral adversarial supervision, thereby supporting natural waveform synthesis. Experiments on LJSpeech, VCTK, and LibriTTS show that FNH-TTS achieves the highest mean MOS on LJSpeech and VCTK and the highest duration-category accuracy on LibriTTS among the compared systems, together with competitive waveform reconstruction and substantially faster vocoder inference. Controlled analyses further show that MoE-DP primarily drives the duration-modeling gains, while the vocoder-side components make complementary contributions to synthesis quality and efficiency.
From: Tian Li [view email]
[v1]
Sat, 16 Aug 2025 10:04:21 UTC (6,179 KB)
[v2]
Tue, 19 Aug 2025 19:48:49 UTC (6,179 KB)
[v3]
Thu, 28 May 2026 14:34:05 UTC (4,112 KB)
[v4]
Wed, 2 Sep 2026 22:34:23 UTC (4,121 KB)
此内容由惯性聚合(RSS阅读器)自动聚合整理,仅供阅读参考。 原文来自 — 版权归原作者所有。