










Abstract:Robust selective auditory attention under multilingual interference is critical for reliable LALM deployment. We introduce MUSI, a cocktail party-inspired controlled diagnostic evaluation of source-grounded spoken-language understanding and reasoning. Each item pairs an English target dialogue with a plausible distractor in English, Spanish, Korean, or Chinese, and evaluates models under (1) single-stream, (2) separation-based, and (3) end-to-end cocktail party settings across controlled SNRs. Across four open-weight and two closed-source LALMs, we find model-dependent language-conditioned variation and heightened vulnerability at adverse SNRs. Errors are dominated by distractor-grounded source confusion, while separation reduces acoustic overlap but often leaves source attribution unresolved. These findings highlight selective auditory attention as an primary capability for reliable LALMs deployment.
From: Heejoon Koo [view email]
[v1]
Sun, 17 May 2026 02:13:58 UTC (383 KB)
[v2]
Thu, 17 Sep 2026 16:11:59 UTC (386 KB)
此内容由惯性聚合(RSS阅读器)自动聚合整理,仅供阅读参考。 原文来自 — 版权归原作者所有。