惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

Know Your Adversary
Know Your Adversary
N
News and Events Feed by Topic
C
CXSECURITY Database RSS Feed - CXSecurity.com
P
Privacy & Cybersecurity Law Blog
P
Privacy International News Feed
Recent Announcements
Recent Announcements
T
Threatpost
D
Darknet – Hacking Tools, Hacker News & Cyber Security
Last Week in AI
Last Week in AI
博客园 - 叶小钗
AWS News Blog
AWS News Blog
C
Cyber Attacks, Cyber Crime and Cyber Security
Cyberwarzone
Cyberwarzone
月光博客
月光博客
The Cloudflare Blog
罗磊的独立博客
P
Palo Alto Networks Blog
酷 壳 – CoolShell
酷 壳 – CoolShell
博客园 - 聂微东
cs.AI updates on arXiv.org
cs.AI updates on arXiv.org
T
Tailwind CSS Blog
PCI Perspectives
PCI Perspectives
Forbes - Security
Forbes - Security
Latest news
Latest news
V
Visual Studio Blog
阮一峰的网络日志
阮一峰的网络日志
奇客Solidot–传递最新科技情报
奇客Solidot–传递最新科技情报
WordPress大学
WordPress大学
让小产品的独立变现更简单 - ezindie.com
让小产品的独立变现更简单 - ezindie.com
有赞技术团队
有赞技术团队
N
News and Events Feed by Topic
爱范儿
爱范儿
L
Lohrmann on Cybersecurity
Cisco Talos Blog
Cisco Talos Blog
I
Intezer
S
Secure Thoughts
Hacker News: Ask HN
Hacker News: Ask HN
C
CERT Recently Published Vulnerability Notes
TaoSecurity Blog
TaoSecurity Blog
博客园 - 司徒正美
Google DeepMind News
Google DeepMind News
cs.CV updates on arXiv.org
cs.CV updates on arXiv.org
Google Online Security Blog
Google Online Security Blog
H
Hacker News: Front Page
Recent Commits to openclaw:main
Recent Commits to openclaw:main
V2EX - 技术
V2EX - 技术
Microsoft Azure Blog
Microsoft Azure Blog
O
OpenAI News
cs.CL updates on arXiv.org
cs.CL updates on arXiv.org
Webroot Blog
Webroot Blog

stat updates on arXiv.org

Improved Baselines with Representation Autoencoders Calibeating for general proper losses: A Bregman divergence approach Dimension-Free Convergence of Discrete Diffusion Models: Adjoint Equations Induce the Right Space XAI and Statistical Analysis for Reliable Intrusion Detection in the UAVIDS-2025 Dataset: From Tree to Hybrid and Tabular DNN Ensembles Reasoning Models Don't Just Think Longer, They Move Differently TabPFN-3: Technical Report Reframing preprocessing selection as model-internal calibration in near-infrared spectroscopy: A large-scale benchmark of operator-adaptive PLS and Ridge models Towards a holistic understanding of Selection Bias for Causal Effect Identification Adaptive Kernel Density Estimation with Pre-training Coreset-Induced Conditional Velocity Flow Matching RISED: A Pre-Deployment Evaluation Framework for High-Stakes AI Decision-Support Systems, with Application to Healthcare ISOMORPH: A Supply Chain Digital Twin for Simulation, Dataset Generation, and Forecasting Benchmarks Yield Curves Dynamics Using Variational Autoencoders Under No-arbitrage Model-based Bootstrap of Controlled Markov Chains Online Learning-to-Defer with Varying Experts Self-Supervised Laplace Approximation for Bayesian Uncertainty Quantification Keeping Score: Efficiency Improvements in Neural Likelihood Surrogate Training via Score-Augmented Loss Functions One-Step Generative Modeling via Wasserstein Gradient Flows Exact Stiefel Optimization for Probabilistic PLS: Closed-Form Updates, Error Bounds, and Calibrated Uncertainty A Composite Activation Function for Learning Stable Binary Representations Adaptive Calibration in Non-Stationary Environments Real vs. Semi-Simulated: Rethinking Evaluation for Treatment Effect Estimation Federated Language Models Under Bandwidth Budgets: Distillation Rates and Conformal Coverage On Variance Reduction in Learning Mean Flows When Attention Beats Fourier: Multi-Scale Transformers for PDE Solving on Irregular Domains A Refined Generalization Analysis for Extreme Multi-class Supervised Contrastive Representation Learning Ensemble Distributionally Robust Bayesian Optimisation The Proxy Presumption: From Semantic Embeddings to Valid Social Measures Modulated learning for private and distributed regression with just a single sample per client device Query-efficient model evaluation using cached responses Order-Agnostic Autoregressive Modelling with Missing Data Grokking or Glitching? How Low-Precision Drives Slingshot Loss Spikes Tuning Derivatives for Causal Fairness in Machine Learning Spherical Flows for Sampling Categorical Data Bayesian Rain Field Reconstruction using Commercial Microwave Links and Diffusion Model Priors Unified Framework of Distributional Regret in Multi-Armed Bandits and Reinforcement Learning Jacobian-Velocity Bounds for Deployment Risk Under Covariate Drift Self-Attention as Transport: Limits of Symmetric Spectral Diagnostics Perturbation is All You Need for Extrapolating Language Models Graph Convolutional Support Vector Regression for Robust Spatiotemporal Forecasting of Urban Air Pollution Segmenting Human-LLM Co-authored Text via Change Point Detection Stochastic Schrödinger Diffusion Models for Pure-State Ensemble Generation Understanding Self-Supervised Learning via Latent Distribution Matching Imbalanced Classification under Capacity Constraints On the Spectral Structure and Objective Equivalence of Orthogonal Multilabel Fisher Discriminants Partially Observed Structural Causal Models Robust and Fast Training via Per-Sample Clipping Efficient Preference Poisoning Attack on Offline RLHF Joint Energy Management and Coordinated AIGC Workload Scheduling for Distributed Data Centers: A Diffusion-Aided Reward Shaping Approach Distributional Causal Mediation via Conditional Generative Modeling A Theory of Saddle Escape in Deep Nonlinear Networks Adaptive Querying with AI Persona Priors Optimal Spatio-Temporal Decoupling for Bayesian Conformal Prediction Wasserstein Distributionally Robust Regret Optimization for Reinforcement Learning from Human Feedback Electricity price forecasting across Norway's five bidding zones in the post-crisis era Adversarial Robustness of NTK Neural Networks Identifiability and Stability of Generative Drifting with Companion-Elliptic Kernel Families A Limit Theory of Foundation Models: A Mathematical Approach to Understanding Emergent Intelligence and Scaling Laws Conditional Score-Based Modeling of Effective Langevin Dynamics Generative Synthetic Data for Causal Inference: Pitfalls, Remedies, and Opportunities Inference of Online Newton Methods with Nesterov's Accelerated Sketching ProEval: Proactive Failure Discovery and Efficient Performance Estimation for Generative AI Evaluation Turtle shell clustering: A mixture approach to discriminative clustering with applications to flow cytometry and other data Score-Repellent Monte Carlo: Toward Efficient Non-Markovian Sampler with Constant Memory in General State Spaces Learning to Emulate Chaos: Adversarial Optimal Transport Regularization Geometric Layer-wise Approximation Rates for Deep Networks S2MAM: Semi-supervised Meta Additive Model for Robust Estimation and Variable Selection Beyond Coefficients: Forecast-Necessity Testing for Interpretable Causal Discovery in Nonlinear Time-Series Models Curiosity-Critic: Cumulative Prediction Error Improvement as a Tractable Intrinsic Reward for World Model Training Knowing When to Quit: A Principled Framework for Dynamic Abstention in LLM Reasoning Enhancing AI and Dynamical Subseasonal Forecasts with Probabilistic Bias Correction MinShap: A Modified Shapley Value Approach for Feature Selection Zeroth-Order Optimization at the Edge of Stability Generative Augmented Inference Estimating Continuous Treatment Effects with Two-Stage Kernel Ridge Regression Rare Event Analysis via Stochastic Optimal Control Adaptive Learning via Off-Model Training and Importance Sampling for Fully Non-Markovian Optimal Stochastic Control. Complete version Beyond Fixed False Discovery Rates: Post-Hoc Conformal Selection with E-Variables Beyond Augmented-Action Surrogates for Multi-Expert Learning-to-Defer Spatio-temporal probabilistic forecast using MMAF-guided learning Conformal Policy Control The Implicit Curriculum: Learning Dynamics in RL with Verifiable Rewards Probabilistic NDVI Forecasting from Sparse Satellite Time Series and Weather Covariates Constrained Policy Optimization with Cantelli-Bounded Value-at-Risk Factorizable joint shift revisited Feature Learning Dynamics in Infinite-Depth Neural Networks Statistically-Guided Meta-Learning for Cross-Deployment Activity Recognition in Distributed Fiber-Optic Sensing DAPS++: Rethinking Diffusion Inverse Problems with Decoupled Posterior Annealing Branching Flows: Discrete, Continuous, and Manifold Flow Matching with Splits and Deletions Manifold Dimension Estimation via Local Graph Structure Neural ARFIMA model for forecasting BRIC exchange rates with long memory Neural Stochastic Differential Equations on Compact State Spaces: Theory, Methods, and Application to Suicide Risk Modeling BOOST: A Data-Driven Framework for the Automated Joint Selection of Kernel and Acquisition Functions in Bayesian Optimization Random Walk Learning and the Pac-Man Attack Random Matrix Theory for Deep Learning: Beyond Eigenvalues of Linear Models Data Balancing Strategies: A Systematic Survey of Resampling and Augmentation Methods Post-Training Augmentation Invariance Optimizing LLM Inference: Fluid-Guided Online Scheduling with Memory Constraints Dataset-Driven Channel Masks in Transformers for Multivariate Time Series Survival of the Cheapest: Cost-Aware Hardware Adaptation for Adversarial Robustness
To Grok Grokking: Provable Grokking in Ridge Regression
[Submitted on 27 Jan 2026 (v1), last revised 16 Jul 2026 (this v · 2026-01-28 · via stat updates on arXiv.org

View PDF HTML (experimental)

Abstract:We study grokking, the onset of generalization long after overfitting, in a classical ridge regression setting. We prove end-to-end grokking results for learning over-parameterized linear regression models using gradient descent with weight decay. Specifically, we prove that the following stages occur: (i) the model overfits the training data early during training; (ii) poor generalization persists long after overfitting has manifested; and (iii) the generalization error eventually becomes arbitrarily small. Moreover, we show, both theoretically and empirically, that grokking can be amplified or eliminated in a principled manner through proper hyperparameter tuning. To the best of our knowledge, these are the first rigorous quantitative bounds on the generalization delay (which we refer to as the "grokking time") in terms of training hyperparameters. Lastly, going beyond the linear setting, we empirically demonstrate that our quantitative bounds also capture the behavior of grokking on non-linear neural networks. Our results suggest that grokking is not an inherent failure mode of deep learning, but rather a consequence of specific training conditions, and thus does not require fundamental changes to the model architecture or learning algorithm to avoid.

Submission history

From: Mingyue Xu [view email]
[v1] Tue, 27 Jan 2026 16:52:04 UTC (1,520 KB)
[v2] Thu, 5 Feb 2026 19:04:09 UTC (1,521 KB)
[v3] Thu, 28 May 2026 21:17:02 UTC (2,128 KB)
[v4] Thu, 16 Jul 2026 15:51:18 UTC (2,128 KB)