























Abstract:We introduce Sparse Concept Anchoring, a method that biases latent space to position a targeted subset of concepts while allowing others to self-organize, using only minimal supervision (labels for <0.1% of examples per anchored concept). Training combines activation normalization, a separation regularizer, and anchor or subspace regularizers that attract rare labeled examples to predefined directions or axis-aligned subspaces. The anchored geometry enables two practical interventions: reversible behavioral steering that projects out a concept's latent component at inference, and permanent removal via targeted weight ablation of anchored dimensions. Experiments on structured autoencoders show selective attenuation of targeted concepts with negligible impact on orthogonal features, and complete elimination with reconstruction error approaching theoretical bounds. Sparse Concept Anchoring therefore provides a practical pathway to interpretable, steerable behavior in learned representations.
| Comments: | 8 pages, 3 figures, 1 table (main text). v2: Renamed sections 3.2, 3.3; Length reduction without substantial changes |
| Subjects: | Machine Learning (cs.LG) |
| ACM classes: | I.2.6 |
| Cite as: | arXiv:2512.12469 [cs.LG] |
| (or arXiv:2512.12469v3 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2512.12469 arXiv-issued DOI via DataCite |
From: Patryk Wielopolski [view email]
[v1]
Sat, 13 Dec 2025 21:43:17 UTC (3,925 KB)
[v2]
Mon, 26 Jan 2026 09:03:00 UTC (3,931 KB)
[v3]
Sat, 25 Apr 2026 18:15:35 UTC (3,925 KB)
此内容由惯性聚合(RSS阅读器)自动聚合整理,仅供阅读参考。 原文来自 — 版权归原作者所有。