

























Abstract:As models grow more capable, humans cannot reliably verify what they say. Scalable steering requires methods that are internal, self-supervised, and transfer out-of-distribution; existing methods satisfy some but not all three. We introduce AntiPaSTO, which separates representations along an antiparallel axis (+1/-1 produce opposite shifts), with coherence constraints preventing collapse. Training uses only two contrasting words inserted into template sentences, with no preference labels. When we use 800 such synthetic pairs on Gemma-3-1B, AntiPaSTO beats prompting baselines by 6.9x Steering F1 on DailyDilemmas and wins on 5 of 6 tested value axes. We also find preliminary evidence that it maintains bidirectional control where prompting triggers refusal.
| Comments: | Code is available at this https URL |
| Subjects: | Machine Learning (cs.LG) |
| Cite as: | arXiv:2601.07473 [cs.LG] |
| (or arXiv:2601.07473v5 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2601.07473 arXiv-issued DOI via DataCite |
From: Michael J. Clark [view email]
[v1]
Mon, 12 Jan 2026 12:27:01 UTC (175 KB)
[v2]
Sat, 17 Jan 2026 22:32:43 UTC (176 KB)
[v3]
Sun, 1 Feb 2026 03:23:44 UTC (312 KB)
[v4]
Sat, 18 Apr 2026 22:38:19 UTC (282 KB)
[v5]
Tue, 12 May 2026 00:51:44 UTC (283 KB)
此内容由惯性聚合(RSS阅读器)自动聚合整理,仅供阅读参考。 原文来自 — 版权归原作者所有。