
























Abstract:While InfoNCE underlies modern contrastive learning, its geometric mechanisms remain under-characterized beyond the canonical alignment--uniformity decomposition. We develop a measure-theoretic framework in which learning evolves representation measures on a fixed embedding manifold. In the large-batch limit, we prove value and gradient consistency, linking the stochastic objective to explicit deterministic energy landscapes and revealing a geometric bifurcation between unimodal and symmetric multimodal regimes. In the unimodal case, the intrinsic energy is strictly convex and admits a unique Gibbs equilibrium, showing that entropy acts as a tie-breaker within the aligned basin. In the multimodal case, the intrinsic geometry becomes cross-coupled and contains a persistent negative symmetric divergence term: each modality's marginal reshapes the effective landscape of the other, allowing strong pairwise alignment to coexist with a persistent modality gap. Controlled synthetic experiments and analyses of pretrained CLIP representations support these predictions. Overall, our results shift the analytical lens from pointwise discrimination to population geometry, showing that pairwise alignment alone is insufficient to control cross-modal marginal structure.
| Comments: | ICML 2026 |
| Subjects: | Machine Learning (cs.LG) |
| Cite as: | arXiv:2601.19597 [cs.LG] |
| (or arXiv:2601.19597v3 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2601.19597 arXiv-issued DOI via DataCite |
From: Yichao Cai [view email]
[v1]
Tue, 27 Jan 2026 13:33:03 UTC (1,427 KB)
[v2]
Mon, 16 Mar 2026 02:05:00 UTC (2,137 KB)
[v3]
Mon, 4 May 2026 12:26:29 UTC (1,922 KB)
此内容由惯性聚合(RSS阅读器)自动聚合整理,仅供阅读参考。 原文来自 — 版权归原作者所有。