






















Abstract:We introduce softpick, a rectified, not sum-to-one, drop-in replacement for softmax in transformer attention mechanisms that eliminates attention sink and massive activations. Our experiments with 340M and 1.8B parameter models demonstrate that softpick achieves 0\% sink rate consistently. The softpick transformers produce hidden states with significantly lower kurtosis and creates sparse attention maps. Quantized models using softpick outperform softmax on standard benchmarks, with a particularly pronounced advantage at lower bit precisions. Our analysis and discussion shows how softpick has the potential to open new possibilities for quantization, low-precision training, sparsity optimization, pruning, and interpretability. Our code: this https URL
| Comments: | Updated to camera-ready version |
| Subjects: | Machine Learning (cs.LG) |
| Cite as: | arXiv:2504.20966 [cs.LG] |
| (or arXiv:2504.20966v4 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2504.20966 arXiv-issued DOI via DataCite |
From: Zayd Muhammad Kawakibi Zuhri [view email]
[v1]
Tue, 29 Apr 2025 17:36:18 UTC (15,254 KB)
[v2]
Fri, 30 May 2025 12:37:29 UTC (10,479 KB)
[v3]
Tue, 13 Jan 2026 11:54:31 UTC (5,728 KB)
[v4]
Fri, 17 Apr 2026 09:56:07 UTC (5,728 KB)
此内容由惯性聚合(RSS阅读器)自动聚合整理,仅供阅读参考。 原文来自 — 版权归原作者所有。