惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

F
Fortinet All Blogs
aimingoo的专栏
aimingoo的专栏
V
Visual Studio Blog
罗磊的独立博客
爱范儿
爱范儿
J
Java Code Geeks
博客园 - 司徒正美
N
Netflix TechBlog - Medium
Microsoft Security Blog
Microsoft Security Blog
美团技术团队
小众软件
小众软件
Google DeepMind News
Google DeepMind News
freeCodeCamp Programming Tutorials: Python, JavaScript, Git & More
V
V2EX
博客园 - 聂微东
云风的 BLOG
云风的 BLOG
WordPress大学
WordPress大学
H
Hackread – Cybersecurity News, Data Breaches, AI and More
Jina AI
Jina AI
Y
Y Combinator Blog
博客园 - 叶小钗
人人都是产品经理
人人都是产品经理
Martin Fowler
Martin Fowler
Vercel News
Vercel News

eess.AS updates on arXiv.org

Dependence on Early and Late Reverberation of Single-Channel Speaker Distance Estimation MIST: Multimodal Interactive Speech-based Tool-calling Conversational Assistants for Smart Homes LiVeAction: a Lightweight, Versatile, and Asymmetric Neural Codec Design for Real-time Operation Weight-Decay Turns Transformer Loss Landscapes Villani: Functional-Analytic Foundations for Optimization and Generalization PairAlign: A Framework for Sequence Tokenization via Self-Alignment with Applications to Audio Tokenization WavCube: Unifying Speech Representation for Understanding and Generation via Semantic-Acoustic Joint Modeling Predictive-Generative Drift Decomposition for Speech Enhancement and Separation Minimizing Modality Gap from the Input Side: Your Speech LLM Can Be a Prosody-Aware Text LLM X-Voice: Enabling Everyone to Speak 30 Languages via Zero-Shot Cross-Lingual Voice Cloning JASTIN: Aligning LLMs for Zero-Shot Audio and Speech Evaluation via Natural Language Instructions Phoneme-Level Deepfake Detection Across Emotional Conditions Using Self-Supervised Embeddings When Audio-Language Models Fail to Leverage Multimodal Context for Dysarthric Speech Recognition Dimensionality-Aware Anomaly Detection in Learned Representations of Self-Supervised Speech Models Mitigating Multimodal LLMs Hallucinations via Relevance Propagation at Inference Time Virtual Speech Therapist: A Clinician-in-the-Loop AI Speech Therapy Agent for Personalized and Supervised Therapy LASE: Language-Adversarial Speaker Encoding for Indic Cross-Script Identity Preservation Towards Improving Speaker Distance Estimation through Generative Impulse Response Augmentation Beyond Decodability: Reconstructing Language Model Representations with an Encoding Probe MMAudioReverbs: Video-Guided Acoustic Modeling for Dereverberation and Room Impulse Response Estimation Alethia: A Foundational Encoder for Voice Deepfakes From Birdsong to Rumbles: Classifying Elephant Calls with Out-of-Species Embeddings Beyond the Baseband: Adaptive Multi-Band Encoding for Full-Spectrum Bioacoustics Classification Predicting Upcoming Stuttering Events from Three-Second Audio: Stratified Evaluation Reveals Severity-Selective Precursors, and the Model Deploys Fully On-Device The False Resonance: A Critical Examination of Emotion Embedding Similarity for Speech Generation Evaluation DiffAnon: Diffusion-based Prosody Control for Voice Anonymization Recurrence-Based Nonlinear Vocal Dynamics as Digital Biomarkers for Depression Detection from Conversational Speech One Voice, Many Tongues: Cross-Lingual Voice Cloning for Scientific Speech Similarity Choice and Negative Scaling in Supervised Contrastive Learning for Deepfake Audio Detection Walking Through Uncertainty: An Empirical Study of Uncertainty Estimation for Audio-Aware Large Language Models Praxy Voice: Voice-Prompt Recovery + BUPS for Commercial-Class Indic TTS from a Frozen Non-Indic Base at Zero Commercial-Training-Data Cost
FNH-TTS: Mixture-of-Experts Duration Modeling for Robust ...
[Submitted on 16 Aug 2025 (v1), last revised 2 Sep 2026 (this ve · 2025-08-16 · via eess.AS updates on arXiv.org

View PDF HTML (experimental)

Abstract:Natural and human-like speech depends on the coordination between prosodic timing and acoustic realization: duration modeling shapes rhythmic structure, while waveform generation determines whether that structure is rendered naturally. In natural speech, duration patterns vary across linguistic contexts and speakers, requiring a TTS system both to capture this variability and to faithfully realize it in the waveform. To address these challenges, we propose FNH-TTS, a VITS-based end-to-end system that jointly improves duration modeling and waveform generation. A mixture-of-experts duration predictor (MoE-DP) uses multiple experts and routing jointly conditioned on linguistic context and speaker information to model diverse duration patterns. For waveform generation, we adopt an inverse short-time Fourier transform (ISTFT)-based generator, providing a more direct and efficient synthesis path. We further employ multi-resolution and sub-band discriminators for fine-grained temporal and spectral adversarial supervision, thereby supporting natural waveform synthesis. Experiments on LJSpeech, VCTK, and LibriTTS show that FNH-TTS achieves the highest mean MOS on LJSpeech and VCTK and the highest duration-category accuracy on LibriTTS among the compared systems, together with competitive waveform reconstruction and substantially faster vocoder inference. Controlled analyses further show that MoE-DP primarily drives the duration-modeling gains, while the vocoder-side components make complementary contributions to synthesis quality and efficiency.

Submission history

From: Tian Li [view email]
[v1] Sat, 16 Aug 2025 10:04:21 UTC (6,179 KB)
[v2] Tue, 19 Aug 2025 19:48:49 UTC (6,179 KB)
[v3] Thu, 28 May 2026 14:34:05 UTC (4,112 KB)
[v4] Wed, 2 Sep 2026 22:34:23 UTC (4,121 KB)