
























Abstract:Gradient-based attribution methods are model-faithful and scalable, but Integrated Gradients (IG) can be brittle because explanations depend on heuristic baselines, straight-line paths, discretization, and saturation. We propose Fisher--Rao Integrated Gradients (FRInGe), which defines both the reference and interpolation schedule in predictive distribution space. FRInGe replaces input baselines with a maximum-entropy predictive reference and follows a Fisher-Rao geodesic on the probability simplex. The corresponding input-space trajectory is realized through the pullback Fisher metric and stabilized by KL and Euclidean trust regions; attributions are obtained by integrating input gradients along this trajectory. Across six ImageNet architectures, FRInGe most clearly improves calibration-oriented attribution metrics, especially MAS scores, while remaining competitive on perturbation AUC and infidelity.
| Subjects: | Machine Learning (cs.LG) |
| Cite as: | arXiv:2605.06404 [cs.LG] |
| (or arXiv:2605.06404v1 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2605.06404 arXiv-issued DOI via DataCite (pending registration) |
From: Gabriele Martino [view email]
[v1]
Thu, 7 May 2026 15:12:15 UTC (42,799 KB)
此内容由惯性聚合(RSS阅读器)自动聚合整理,仅供阅读参考。 原文来自 — 版权归原作者所有。