







Abstract:This work presents a framework for compressing self-supervised models for speaker diarization through structured pruning guided by knowledge distillation. We investigate pruning objectives that target reducing both model parameters and computational complexity, where knowledge distillation enables compact models and pruning removes unnecessary parameters to improve hardware efficiency. We further analyze alternative pruning strategies, showing that a simple overall pruning approach provides the best balance between efficiency and accuracy. Compared to the original unpruned model, our method achieves up to 80% model size reduction and 4x faster inference without performance degradation. Comprehensive experiments across eight public diarization datasets demonstrate that the pruned models consistently match or surpass the performance of their uncompressed counterparts. Furthermore, we show strong outof-domain generalization on the CHiME-6 dataset, achieving accuracy comparable to the top systems in the CHiME-7 challenge without any domain adaptation. These results highlight that structured pruning, when guided by distillation, can yield efficient and generalizable diarization systems suitable for realworld applications.
From: Jiangyu Han [view email]
[v1]
Mon, 23 Jun 2025 13:29:51 UTC (410 KB)
[v2]
Wed, 19 Nov 2025 14:45:23 UTC (392 KB)
[v3]
Wed, 16 Sep 2026 09:05:31 UTC (394 KB)
此内容由惯性聚合(RSS阅读器)自动聚合整理,仅供阅读参考。 原文来自 — 版权归原作者所有。