





















Abstract:Chronic diseases often progress differently across patients. Rather than randomly varying, there are typically a small number of subtypes for how a disease progresses across patients. To capture this structured heterogeneity, the Subtype and Stage Inference Event-Based Model (SuStaIn) estimates the number of subtypes, the order of disease progression for each subtype, and assigns each patient to a subtype from primarily cross-sectional data. It has been widely applied to uncover the subtypes of many diseases and inform our understanding of them. But how robust is its performance? In this paper, we develop a principled Bayesian subtype variant of the event-based model (BEBMS) and compare its performance to SuStaIn in a variety of synthetic data experiments with varied levels of model misspecification. BEBMS substantially outperforms SuStaIn across ordering, staging, and subtype assignment tasks. Further, we apply BEBMS and SuStaIn to a real-world Alzheimer's data set. We find BEBMS has results that are more consistent with the scientific consensus of Alzheimer's disease progression than SuStaIn.
| Comments: | 32 pages; machine learning for health symposium (2025); Proceedings of the 5th Machine Learning for Health Symposium in PMLR |
| Subjects: | Machine Learning (cs.LG); Methodology (stat.ME) |
| Cite as: | arXiv:2512.03467 [cs.LG] |
| (or arXiv:2512.03467v2 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2512.03467 arXiv-issued DOI via DataCite |
|
| Journal reference: | Proceedings of Machine Learning Research (PMLR), vol. 297: Machine Learning for Health (ML4H), 2025 |
From: Hongtao Hao [view email]
[v1]
Wed, 3 Dec 2025 05:45:16 UTC (511 KB)
[v2]
Tue, 21 Apr 2026 17:13:47 UTC (488 KB)
此内容由惯性聚合(RSS阅读器)自动聚合整理,仅供阅读参考。 原文来自 — 版权归原作者所有。