










Abstract:Deep learning-based Sound Event Localization and Detection (SELD) systems suffer severe performance degradation in real-world, long-tailed acoustic environments. Standard continuous regression objectives heavily bias learning toward frequent classes, causing rare events to be systematically under-recognized, an optimization bottleneck we term detection timidity. To overcome this, we propose MAGENTA (Magnitude And Geometry-ENhanced Training Approach), an architecture-agnostic loss framework that geometrically decomposes the regression error into orthogonal radial (activity) and angular (localization) components. Unlike standard methods that rely on static frequency weights, MAGENTA incorporates an intrinsic, difficulty-driven annealing mechanism. By decoupling the objective to independently modulate active detection and inactive suppression, the system can adaptively boost recall for difficult tail classes while modulating inactive penalties to prevent spurious rare-event detections. Evaluations on the STARSS23 dataset demonstrate that MAGENTA yields a 20.5% relative reduction in the aggregated SELD error, effectively recovering tail class performance without compromising head class precision. Code is available at: this https URL
From: Jun Wei Yeow [view email]
[v1]
Fri, 19 Sep 2025 05:00:20 UTC (1,757 KB)
[v2]
Tue, 4 Aug 2026 07:34:01 UTC (1,361 KB)
此内容由惯性聚合(RSS阅读器)自动聚合整理,仅供阅读参考。 原文来自 — 版权归原作者所有。