



























Abstract:We revisit the classical broken sample problem: Two samples of i.i.d.\ data points ${\mathbf{X}}=\{X_{1},\ldots , X_{n}\}$ and ${\mathbf{Y}}=\{Y_{1},\ldots ,Y_{m}\}$ are observed without correspondence with $m\leq n$. Under the null hypothesis, ${\mathbf{X}}$ and ${\mathbf{Y}}$ are independent. Under the alternative hypothesis, ${\mathbf{Y}}$ is correlated with a random subsample of ${\mathbf{X}}$, in the sense that $(X_{\pi (i)},Y_{i})$'s are drawn independently from some bivariate distribution for some latent injection $\pi :[m] \to [n]$. Originally introduced by DeGroot, Feder, and Goel to model matching records in census data, this problem has recently gained renewed interest due to its applications in data de-anonymization, data integration, and target tracking. Despite extensive research over the past decades, determining the precise detection threshold has remained an open problem even for equal sample sizes ($m=n$). Assuming $m$ and $n$ grow proportionally, we show that the sharp threshold is given by a spectral and an $L_{2}$ condition of the likelihood ratio operator, resolving a conjecture of Bai and Hsing in the positive. These results are extended to high dimensions and settle the sharp detection thresholds for Gaussian and Bernoulli models.
From: Simiao Jiao [view email]
[v1]
Tue, 18 Mar 2025 18:15:13 UTC (105 KB)
[v2]
Sat, 4 Jul 2026 00:35:38 UTC (108 KB)
此内容由惯性聚合(RSS阅读器)自动聚合整理,仅供阅读参考。 原文来自 — 版权归原作者所有。