





















Abstract:Several methods in statistics and machine learning target a probability distribution for which an entropy-regularised variational objective is minimised. This increased flexibility introduces a computational challenge, as one loses access to an explicit unnormalised density for the target. To mitigate this difficulty, we introduce a novel measure of suboptimality called 'gradient discrepancy', and in particular a kernel gradient discrepancy (KGD) that can be explicitly computed. In the Bayesian statistics context, KGD coincides with the kernel Stein discrepancy (KSD), and we obtain a novel characterisation of KSD as measuring the size of a variational gradient. Outside this familiar setting, KGD enables novel sampling algorithms to be developed and compared, even when unnormalised densities cannot be obtained. To illustrate this point several novel algorithms are proposed and studied, including a natural generalisation of Stein variational gradient descent, with applications to mean-field neural networks and predictively oriented posteriors presented. On the theoretical side, our principal contribution is to establish sufficient conditions for desirable properties of KGD, such as continuity and convergence control.
From: Chris Oates [view email]
[v1]
Fri, 12 Sep 2025 16:38:41 UTC (879 KB)
[v2]
Fri, 17 Oct 2025 11:51:16 UTC (2,294 KB)
[v3]
Tue, 16 Dec 2025 10:38:09 UTC (2,479 KB)
[v4]
Thu, 9 Jul 2026 19:22:33 UTC (2,755 KB)
此内容由惯性聚合(RSS阅读器)自动聚合整理,仅供阅读参考。 原文来自 — 版权归原作者所有。