













Abstract:Children's automatic speech recognition (ASR) remains challenging because child speech differs from adult speech and varies substantially across developmental stages. While adapter tuning provides a promising way to adapt large pretrained ASR models to children's speech, a single shared child adapter may not fully capture age-dependent variation. In this work, we present one of the first systematic studies of age-aware adapter tuning for child ASR, focusing on speech from children aged 3-12 and older. We propose age-specialized adapters trained separately for different age groups and compare them with a unified age-conditioned FiLM adapter. With ground-truth age routing, age-specialized adapters improve over a strong shared child adapter baseline from 12.5% to 12.3% overall word error rate (WER) and from 16.4% to 16.1% macro-age WER, while consistently improving WER across all four known-age groups. We further show that predicted-age routing remains close to ground-truth routing, achieving 12.3% overall WER and 16.3% macro-age WER without ground-truth age labels at inference. In contrast, unified FiLM conditioning does not consistently improve over the shared child adapter, indicating that a single unified adapter may be insufficient to capture developmental variation in child speech.
From: Jialu Li [view email]
[v1]
Wed, 3 Jun 2026 21:02:28 UTC (107 KB)
[v2]
Mon, 14 Sep 2026 17:24:25 UTC (113 KB)
此内容由惯性聚合(RSS阅读器)自动聚合整理,仅供阅读参考。 原文来自 — 版权归原作者所有。