











Abstract:Training-free anomalous sound detection (ASD) based on pre-trained audio embedding models has recently garnered significant attention, as it enables the detection of anomalous sounds using only normal reference data without task-specific model training or fine-tuning. However, existing embedding-based approaches almost exclusively rely on temporal mean pooling, leaving temporal pooling in training-free ASD largely unexplored. In this paper, we present the first systematic evaluation of temporal pooling strategies for training-free ASD with pre-trained audio embeddings. We propose relative deviation pooling (RDP), an adaptive pooling method that assigns larger weights to embeddings with stronger temporal deviations, investigate feature-wise non-linear aggregation using generalized mean (GeM) pooling, and examine a hybrid combination of both strategies. Experiments on five benchmark datasets demonstrate that the proposed pooling strategies consistently outperform mean pooling and achieve state-of-the-art performance for training-free ASD, including results that surpass previously reported trained systems and ensembles on the DCASE2025 ASD dataset.
From: Kevin Wilkinghoff [view email]
[v1]
Wed, 4 Mar 2026 21:07:22 UTC (116 KB)
[v2]
Sun, 21 Jun 2026 15:47:17 UTC (119 KB)
[v3]
Thu, 6 Aug 2026 06:12:27 UTC (118 KB)
此内容由惯性聚合(RSS阅读器)自动聚合整理,仅供阅读参考。 原文来自 — 版权归原作者所有。