



























Abstract:We study the problem of understanding where two populations differ within a feature space, which we formalize in the concept of a differential subgroup: a subset of individuals from both populations who, despite sharing similar characteristics, exhibit exceptional differences in a target outcome. Differential subgroups reveal the regions of the feature space where population-level gaps are most pronounced and can help practitioners identify the covariate combinations that are structurally responsible for these differences, e.g.~in clinical analysis, model diagnostics, or treatment-effect studies. We introduce a general optimization objective for discovering differential subgroups and establish conditions under which the resulting subgroups admit a causal interpretation of population differences. We propose DiffSub, a gradient-based approach that discovers interpretable differential subgroups in tabular data. Across synthetic benchmarks, medical case studies, model-error analyses, and treatment-effect settings, DiffSub identifies informative subgroups that reveal where population differences arise and why.
| Subjects: | Machine Learning (cs.LG) |
| Cite as: | arXiv:2604.27741 [cs.LG] |
| (or arXiv:2604.27741v1 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2604.27741 arXiv-issued DOI via DataCite (pending registration) |
From: Sascha Xu [view email]
[v1]
Thu, 30 Apr 2026 11:31:42 UTC (5,036 KB)
此内容由惯性聚合(RSS阅读器)自动聚合整理,仅供阅读参考。 原文来自 — 版权归原作者所有。