


























Abstract:Self-supervised learning, in the context of foundation model training, is a powerful pre-training method for learning feature representations without labels, which often capture generic underlying semantics from the data and can later be fine-tuned for downstream tasks. In this work, we introduce jBOT, a pre-training method based on self-distillation for jet data from the CERN Large Hadron Collider, which combines local particle-level distillation with global jet-level distillation to learn jet representations that support downstream tasks such as anomaly detection and classification. We observe that pre-training on unlabeled jets leads to emergent semantic class clustering in the representation space. The clustering in the frozen embedding, when pre-trained on background jets only, enables anomaly detection via simple distance-based metrics, and the learned embedding can be fine-tuned for classification with improved performance compared to supervised models trained from scratch.
| Comments: | Under review |
| Subjects: | Machine Learning (cs.LG); High Energy Physics - Experiment (hep-ex) |
| Cite as: | arXiv:2601.11719 [cs.LG] |
| (or arXiv:2601.11719v3 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2601.11719 arXiv-issued DOI via DataCite |
From: Ho Fung Tsoi [view email]
[v1]
Fri, 16 Jan 2026 19:12:13 UTC (6,553 KB)
[v2]
Wed, 21 Jan 2026 04:25:59 UTC (6,554 KB)
[v3]
Fri, 24 Apr 2026 00:26:44 UTC (6,557 KB)
此内容由惯性聚合(RSS阅读器)自动聚合整理,仅供阅读参考。 原文来自 — 版权归原作者所有。