






















Abstract:Effective model selection is critical in symbolic regression (SR) to identify mathematical expressions that balance accuracy and complexity, and have low expected error on unseen data. Many modern implementations of genetic programming (GP) for SR generate a set of Pareto optimal candidate solutions, but reliable automatic selection of solutions that generalize well remains an open issue. Current literature offers various information-theoretic and Bayesian approaches, yet comprehensive comparisons of their performance across different data regimes are limited. This study presents a systematic empirical comparison of widely used selection criteria: the Akaike information criterion (AIC), the corrected AIC (AICc), the Bayesian information criterion (BIC), minimum description length (MDL), as well as Efron's bootstrap estimate for the in-sample prediction error on seven synthetic datasets with Gaussian noise. We rank candidate expressions generated by perturbing ground-truth functions to assess generalization error and selection probability of the ground-truth expression. Our findings reveal that MDL consistently identifies models with the lowest test error and the shortest length across most datasets. While no single criterion dominates all results, MDL and BIC produced the highest probability of selecting the ground-truth expressions.
| Subjects: | Machine Learning (cs.LG) |
| Cite as: | arXiv:2605.11233 [cs.LG] |
| (or arXiv:2605.11233v1 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2605.11233 arXiv-issued DOI via DataCite (pending registration) |
|
| Related DOI: | https://doi.org/10.1145/3795095.3805157
DOI(s) linking to related resources |
From: Ali Soltani [view email]
[v1]
Mon, 11 May 2026 20:47:50 UTC (187 KB)
此内容由惯性聚合(RSS阅读器)自动聚合整理,仅供阅读参考。 原文来自 — 版权归原作者所有。