

























Abstract:Vision-Language Models (VLMs) are typically deterministic in nature and lack intrinsic mechanisms to quantify epistemic uncertainty, which reflects the model's lack of knowledge or ignorance of its own representations. We theoretically motivate negative log-density of an embedding as a proxy for the epistemic uncertainty, where low-density regions signify model ignorance. The proposed method REPVLM computes the probability density on the hyperspherical manifold of the VLM embeddings using Riemannian Flow Matching. We empirically demonstrate that REPVLM achieves near-perfect correlation between uncertainty and prediction error, significantly outperforming existing baselines. Beyond classification, we also demonstrate that the model also provides a scalable metric for out-of-distribution detection and automated data curation.
| Subjects: | Machine Learning (cs.LG) |
| Cite as: | arXiv:2601.21662 [cs.LG] |
| (or arXiv:2601.21662v2 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2601.21662 arXiv-issued DOI via DataCite |
|
| Journal reference: | Forty-Third International Conference on Machine Learning, 2026 |
From: Li Ju [view email]
[v1]
Thu, 29 Jan 2026 12:58:42 UTC (2,062 KB)
[v2]
Wed, 20 May 2026 13:03:49 UTC (2,020 KB)
此内容由惯性聚合(RSS阅读器)自动聚合整理,仅供阅读参考。 原文来自 — 版权归原作者所有。