























Abstract:Despite the burgeoning body of work on distribution shifts, provenance shift-where the relationship between data source and label changes at deployment-remains poorly understood and under-addressed. In this paper, we establish a formal connection between provenance shift, counterfactual invariance, and invariant learning to derive a learning objective for robustness. We then introduce \textsc{DeconDTN-Toolkit}, a specialized evaluation and remediation suite designed to simulate provenance shifts of varying degrees while maintaining the training protocol and the infrastructure of existing benchmarks. We reveal the vulnerability of Empirical Risk Minimization under provenance shift, introduce a robust out-of-distribution performance indicator, and conduct a comprehensive evaluation on existing algorithms. Our work provides both the theoretical grounding and the practical tools necessary to characterize the problem of confounding by provenance, and implementations of methods to mitigate it.
| Comments: | Accepted to CHIL 2026 |
| Subjects: | Machine Learning (cs.LG) |
| Cite as: | arXiv:2605.11237 [cs.LG] |
| (or arXiv:2605.11237v1 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2605.11237 arXiv-issued DOI via DataCite (pending registration) |
From: Yongsen Tan [view email]
[v1]
Mon, 11 May 2026 20:54:42 UTC (769 KB)
此内容由惯性聚合(RSS阅读器)自动聚合整理,仅供阅读参考。 原文来自 — 版权归原作者所有。