



















Abstract:This note shows that no self-attention layer post-processed by a rational function can sign-represent the parity function unless the product of the number of heads and the degree of the post-processing function grows linearly with the input length. Combining this lower bound with rational approximation of ReLU networks yields a margin-dependent extension for self-attention layers post-processed by ReLU networks.
| Subjects: | Machine Learning (cs.LG) |
| Cite as: | arXiv:2605.12171 [cs.LG] |
| (or arXiv:2605.12171v1 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2605.12171 arXiv-issued DOI via DataCite (pending registration) |
From: Daniel Hsu [view email]
[v1]
Tue, 12 May 2026 14:17:48 UTC (6 KB)
此内容由惯性聚合(RSS阅读器)自动聚合整理,仅供阅读参考。 原文来自 — 版权归原作者所有。