惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

博客园_首页
B
Blog
V
V2EX
T
Tailwind CSS Blog
Hugging Face - Blog
Hugging Face - Blog
博客园 - 【当耐特】
博客园 - 聂微东
博客园 - 叶小钗
博客园 - 三生石上(FineUI控件)
The Cloudflare Blog
J
Java Code Geeks
H
Help Net Security
雷峰网
雷峰网
Apple Machine Learning Research
Apple Machine Learning Research
H
Hackread – Cybersecurity News, Data Breaches, AI and More
Engineering at Meta
Engineering at Meta
F
Fortinet All Blogs
Martin Fowler
Martin Fowler
D
Docker
L
LangChain Blog
人人都是产品经理
人人都是产品经理
爱范儿
爱范儿
WordPress大学
WordPress大学
V
Visual Studio Blog

math updates on arXiv.org

Coupling-Robust Accuracy in Multiphysics Physics Informed Neural Networks via Kronecker-Preconditioned Optimization Non-normal spectral signatures of instability in neural network training dynamics Optimization of randomized neural networks for transfer operator approximation Selective Ambulance Dispatch Under Contextual Travel-Time Uncertainty LLAMA LIMA: A Living Meta-Analysis on the Effects of Generative AI on Learning Mathematics Neural Flow Operators can Approximate any Operator: Abstract Frameworks and Universal Approximations LLMs as Noisy Channels: A Shannon Perspective on Model Capacity and Scaling Laws On the Stability of Spherical Hellinger-Kantorovich Flows and Their Implications for Differential Privacy Training-Free Looped Transformers Move on Muon : A Hamiltonian probability gradient flow perspective of Muon optimizer Entrywise Error Bounds for Spectral Ranking with Semi-Random Adversaries Asymmetric Scaling Laws from Sparse Features Is Dimensionality a Barrier for Retrieval Models? RA-DCA: A Randomized Active-Set DCA for Directional Stationarity in Max-Structured DC Programs Commutator-Induced Uncertainty in VAEs Weisfeiler-Leman Is Incomplete on Simple Spectrum Graphs, so Canonicalize Them Sparse In-Network Learning via Shortest-Path Backpropagation and Finite-Rate Gating Instance-Optimal Estimation with Multiple LLM Judges on a Budget Entropy Equivalence Testing Expand More, Shrink Less: Shaping Effective-Rank Dynamics for Dense Scaling in Recommendation Any-Dimensional Invariant Universality Operationalizing Individual Fairness via Gradient Descent and Bradley-Terry Models Anytime Training with Schedule-Free Spectral Optimization Diffusion-based Denoising Beats Vanilla Score Matching in Parameter Estimation: A Theoretical Explanation Resilience Characterization of AI-Native Wireless Receivers via Persistent Homology The General Theory of Localization Methods Group-Algebraic Tensors: Provably-optimal Equivariant Learning and Physical Symmetry Discovery General Lower Bounds for Differentially Private Federated Learning with Arbitrary Public-Transcript Interactions PilotWiMAE: Pilot-Native Representation Learning for Wireless Channels Proximal basin hopping: global optimization with guarantees
Retention Profiles and KL Contraction Bounds in Finite Ma...
[Submitted on 25 Jun 2026] · 2026-06-26 · via math updates on arXiv.org

View PDF HTML (experimental)

Abstract:We study Kullback-Leibler (KL) contraction in finite Markov chains through a row-wise perspective. Evaluating the SDPI ratio at point masses yields a state-indexed retention profile $r(x)=D_{\mathrm{KL}}(P(x,\cdot)\|\pi)/\log(1/\pi(x))$ and a localization ratio $L(P)=\bar r_\pi/M\in[0,1]$ (with $M=\max_x r(x)$, $\bar r_\pi=\mathbb{E}_\pi r$) that distinguishes localized from global contraction obstructions. Our main contributions are (i) a convexity-gap identity showing that the gap between the row-averaged divergence and $D_{\mathrm{KL}}(\mu P\|\pi)$ equals the mutual information $I_\mu(X;Y)$, and a derived decomposition of the contraction ratio into entropy inflation and a mutual-information penalty; (ii) a Cheeger-type lower bound on $M$, tying the bottleneck geometry of $P$ directly to the row-retention profile; (iii) an explicit construction proving that $L(P_n)\to 0$ does not force $\eta_{\mathrm{KL}}(P_n)/M_n\to 1$, identifying cardinality of high-retention states (not their $\pi$-mass) as the decisive quantity. Alongside these, we record structural consequences, optimal Markov/reverse-Markov tail bounds for $r$, a Bhatia-Davis variance bound, two-sided spectral bounds with an explicit cubic correction, a KL/Pinsker mixing-time bound, and tensorization for product chains. We further show that $L(P)$ is structurally decoupled from the spectral gap, the Cheeger constant, and the mixing time: every vertex-transitive chain satisfies $L(P)=1$ regardless of its mixing speed, and the empirical rank correlations between $L(P)$ and these classical invariants on a diverse but limited test suite are essentially zero. The numerical experiments are exploratory and not used as evidence for a universal classification theorem.

Submission history

From: Saurav Jadhav [view email]
[v1] Thu, 25 Jun 2026 14:15:59 UTC (32 KB)