









Abstract:Room Impulse Responses (RIRs) enable realistic acoustic simulation, with applications ranging from multimedia production to speech data augmentation. However, acquiring high-quality real-world RIRs is labor-intensive, and data scarcity remains a challenge for data-driven RIR generation approaches. In this paper, we propose a novel approach to RIR generation by adapting a pre-trained text-to-audio model, demonstrating for the first time that large-scale generative audio priors can be effectively leveraged for this task. To address the lack of text-RIR paired data, we utilize a labeling pipeline leveraging vision-language models to extract acoustic descriptions from existing image-RIR datasets. We introduce an in-context learning strategy to accommodate free-form user prompts during inference. Evaluations including a subjective listening test demonstrate that our model generates plausible RIRs with substantially less training data. Audio examples are available on our demo website.
From: Kirak Kim [view email]
[v1]
Tue, 10 Mar 2026 14:17:42 UTC (421 KB)
[v2]
Sat, 9 May 2026 21:30:25 UTC (132 KB)
[v3]
Tue, 12 May 2026 18:20:17 UTC (132 KB)
[v4]
Mon, 27 Jul 2026 06:37:09 UTC (133 KB)
此内容由惯性聚合(RSS阅读器)自动聚合整理,仅供阅读参考。 原文来自 — 版权归原作者所有。