惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

钛媒体:引领未来商业与生活新知
钛媒体:引领未来商业与生活新知
Help Net Security
Help Net Security
Threat Intelligence Blog | Flashpoint
Threat Intelligence Blog | Flashpoint
T
Threat Research - Cisco Blogs
T
The Exploit Database - CXSecurity.com
P
Privacy International News Feed
T
Threatpost
T
Tor Project blog
AWS News Blog
AWS News Blog
S
Schneier on Security
Cyberwarzone
Cyberwarzone
The Hacker News
The Hacker News
Scott Helme
Scott Helme
C
Cybersecurity and Infrastructure Security Agency CISA
Application and Cybersecurity Blog
Application and Cybersecurity Blog
Cyber Security Advisories - MS-ISAC
Cyber Security Advisories - MS-ISAC
P
Palo Alto Networks Blog
P
Proofpoint News Feed
Vercel News
Vercel News
Recent Commits to openclaw:main
Recent Commits to openclaw:main
V
V2EX
腾讯CDC
C
CERT Recently Published Vulnerability Notes
www.infosecurity-magazine.com
www.infosecurity-magazine.com
V2EX - 技术
V2EX - 技术
C
Cyber Attacks, Cyber Crime and Cyber Security
MyScale Blog
MyScale Blog
博客园 - 三生石上(FineUI控件)
有赞技术团队
有赞技术团队
D
Docker
Security Latest
Security Latest
云风的 BLOG
云风的 BLOG
G
Google Developers Blog
Know Your Adversary
Know Your Adversary
宝玉的分享
宝玉的分享
爱范儿
爱范儿
Simon Willison's Weblog
Simon Willison's Weblog
N
News | PayPal Newsroom
Recent Announcements
Recent Announcements
小众软件
小众软件
Project Zero
Project Zero
SecWiki News
SecWiki News
Microsoft Azure Blog
Microsoft Azure Blog
月光博客
月光博客
Cloudbric
Cloudbric
博客园 - Franky
Forbes - Security
Forbes - Security
C
Cisco Blogs
Webroot Blog
Webroot Blog
H
Help Net Security

stat updates on arXiv.org

Variance Reduction for Expectations with Diffusion Teachers Everywhere Valid Bounds on False Discovery Proportions in Conformal Inference Decision-Path Patterns as Tree Reliability Signals: Path-based Adaptive Weighting for Random Forest Classification The General Theory of Localization Methods CASCADE Conformal Prediction: Uncertainty-Adaptive Prediction Intervals for Two-Stage Clinical Decision Support Symmetrization of Loss Functions for Robust Training of Neural Networks in the Presence of Noisy Labels Tail Annealing for Heavy-Tailed Flow Matching Variance-Reduced Manifold Sampling via Polynomial-Maximization Density Estimation Latent Laplace Diffusion for Irregular Multivariate Time Series Precision Physical Activity Prescription via Reinforcement Learning for Functional Actions Reducing Diffusion Model Memorization with Higher Order Langevin Dynamics Provably Data-driven Lagrangian Relaxation for Mixed Integer Linear Programming Can Adaptive Gradient Methods Converge under Heavy-Tailed Noise? A Case Study of AdaGrad Shallow ReLU$^s$ Networks in $L^p$-Type and Sobolev Spaces: Approximation and Path-Norm Controlled Generalization Markov Chain Decoders Overcome the Heavy-Tail Limitations of Lipschitz Generative Models On Stability and Decomposition of Sample Quantiles under Heavy-Tailed Distributions Improved Baselines with Representation Autoencoders Symmetry-Compatible Principle for Optimizer Design: Embeddings, LM Heads, SwiGLU MLPs, and MoE Routers Feature Learning in Linear-Width Two-Layer Networks: Two vs. One Step of Gradient Descent A Two-Parameter Weibull Framework for Diagnosing Transformer Weight Distributions Dimension-Free Convergence of Discrete Diffusion Models: Adjoint Equations Induce the Right Space Sample-efficient inductive matrix completion with noise and inexact side-information Multi-task Linear Regression without Eigenvalue Lower Bounds: Adaptivity, Robustness, and Safety Reasoning Models Don't Just Think Longer, They Move Differently TabPFN-3: Technical Report Reframing preprocessing selection as model-internal calibration in near-infrared spectroscopy: A large-scale benchmark of operator-adaptive PLS and Ridge models Towards a holistic understanding of Selection Bias for Causal Effect Identification Adaptive Kernel Density Estimation with Pre-training Coreset-Induced Conditional Velocity Flow Matching RISED: A Pre-Deployment Evaluation Framework for High-Stakes AI Decision-Support Systems, with Application to Healthcare ISOMORPH: A Supply Chain Digital Twin for Simulation, Dataset Generation, and Forecasting Benchmarks Yield Curves Dynamics Using Variational Autoencoders Under No-arbitrage Online Learning-to-Defer with Varying Experts Self-Supervised Laplace Approximation for Bayesian Uncertainty Quantification Keeping Score: Efficiency Improvements in Neural Likelihood Surrogate Training via Score-Augmented Loss Functions One-Step Generative Modeling via Wasserstein Gradient Flows Exact Stiefel Optimization for Probabilistic PLS: Closed-Form Updates, Error Bounds, and Calibrated Uncertainty A Composite Activation Function for Learning Stable Binary Representations Adaptive Calibration in Non-Stationary Environments Real vs. Semi-Simulated: Rethinking Evaluation for Treatment Effect Estimation Federated Language Models Under Bandwidth Budgets: Distillation Rates and Conformal Coverage When Attention Beats Fourier: Multi-Scale Transformers for PDE Solving on Irregular Domains A Refined Generalization Analysis for Extreme Multi-class Supervised Contrastive Representation Learning Ensemble Distributionally Robust Bayesian Optimisation Modulated learning for private and distributed regression with just a single sample per client device Query-efficient model evaluation using cached responses Order-Agnostic Autoregressive Modelling with Missing Data Grokking or Glitching? How Low-Precision Drives Slingshot Loss Spikes Spherical Flows for Sampling Categorical Data Bayesian Rain Field Reconstruction using Commercial Microwave Links and Diffusion Model Priors Unified Framework of Distributional Regret in Multi-Armed Bandits and Reinforcement Learning Jacobian-Velocity Bounds for Deployment Risk Under Covariate Drift Self-Attention as Transport: Limits of Symmetric Spectral Diagnostics Graph Convolutional Support Vector Regression for Robust Spatiotemporal Forecasting of Urban Air Pollution Stochastic Schrödinger Diffusion Models for Pure-State Ensemble Generation Understanding Self-Supervised Learning via Latent Distribution Matching Imbalanced Classification under Capacity Constraints On the Spectral Structure and Objective Equivalence of Orthogonal Multilabel Fisher Discriminants Robust and Fast Training via Per-Sample Clipping Efficient Preference Poisoning Attack on Offline RLHF Distributional Causal Mediation via Conditional Generative Modeling A Theory of Saddle Escape in Deep Nonlinear Networks Adaptive Querying with AI Persona Priors Optimal Spatio-Temporal Decoupling for Bayesian Conformal Prediction Electricity price forecasting across Norway's five bidding zones in the post-crisis era Adversarial Robustness of NTK Neural Networks Identifiability and Stability of Generative Drifting with Companion-Elliptic Kernel Families A Limit Theory of Foundation Models: A Mathematical Approach to Understanding Emergent Intelligence and Scaling Laws Conditional Score-Based Modeling of Effective Langevin Dynamics Inference of Online Newton Methods with Nesterov's Accelerated Sketching ProEval: Proactive Failure Discovery and Efficient Performance Estimation for Generative AI Evaluation Score-Repellent Monte Carlo: Toward Efficient Non-Markovian Sampler with Constant Memory in General State Spaces Learning to Emulate Chaos: Adversarial Optimal Transport Regularization Geometric Layer-wise Approximation Rates for Deep Networks S2MAM: Semi-supervised Meta Additive Model for Robust Estimation and Variable Selection Beyond Coefficients: Forecast-Necessity Testing for Interpretable Causal Discovery in Nonlinear Time-Series Models Curiosity-Critic: Cumulative Prediction Error Improvement as a Tractable Intrinsic Reward for World Model Training Knowing When to Quit: A Principled Framework for Dynamic Abstention in LLM Reasoning Generative Augmented Inference Estimating Continuous Treatment Effects with Two-Stage Kernel Ridge Regression Rare Event Analysis via Stochastic Optimal Control Adaptive Learning via Off-Model Training and Importance Sampling for Fully Non-Markovian Optimal Stochastic Control. Complete version Beyond Augmented-Action Surrogates for Multi-Expert Learning-to-Defer Spatio-temporal probabilistic forecast using MMAF-guided learning The Implicit Curriculum: Learning Dynamics in RL with Verifiable Rewards Probabilistic NDVI Forecasting from Sparse Satellite Time Series and Weather Covariates Constrained Policy Optimization with Cantelli-Bounded Value-at-Risk Feature Learning Dynamics in Infinite-Depth Neural Networks Statistically-Guided Meta-Learning for Cross-Deployment Activity Recognition in Distributed Fiber-Optic Sensing DAPS++: Rethinking Diffusion Inverse Problems with Decoupled Posterior Annealing Branching Flows: Discrete, Continuous, and Manifold Flow Matching with Splits and Deletions Neural ARFIMA model for forecasting BRIC exchange rates with long memory Neural Stochastic Differential Equations on Compact State Spaces: Theory, Methods, and Application to Suicide Risk Modeling BOOST: A Data-Driven Framework for the Automated Joint Selection of Kernel and Acquisition Functions in Bayesian Optimization Random Walk Learning and the Pac-Man Attack Random Matrix Theory for Deep Learning: Beyond Eigenvalues of Linear Models Post-Training Augmentation Invariance Optimizing LLM Inference: Fluid-Guided Online Scheduling with Memory Constraints Program Evaluation with Remotely Sensed Outcomes Dataset-Driven Channel Masks in Transformers for Multivariate Time Series
LLMs on Tabular Data with Limited Semantics: Evidence from Industrial Car Retrofit Prediction
Aina Vila Pons, Ioannis Tzachristas, Constantinos Antoniou · 2026-06-13 · via stat updates on arXiv.org

Industrial retrofit planning depends on structured operational data rather than free text: planners must estimate whether a newly registered prototype will require a retrofit, which retrofit package it will need, and how long the work will take. We study an industrial dataset linking a prototype-registration system (284,271 vehicles) with a retrofit-management system (48,716 cleaned visits), and compare strong tabular machine learning baselines with three LLM-based strategies on row-serialized inputs: embedding features (Amazon Titan), direct prompted classification (Claude Sonnet 4), and an ML+LLM stacking approach. Across binary occurrence prediction, 15-way retrofit-type classification, per-visit duration regression, and an aggregated monthly benchmark, classical tree ensembles remain the strongest standalone models. However, the LLM results reveal a consistent pattern: embeddings remain useful on tables (binary AUC = 0.982), direct prompting collapses once semantic signal is stripped by hashing (binary AUC = 0.500; multiclass weighted F1 = 0.018), and hybrid stacking yields the best manually built multiclass model (weighted F1 = 0.626). On the monthly benchmark, lag-based machine learning outperforms time-series foundation models, though Chronos-small remains competitive in zero-shot forecasting. The results suggest that on privacy-constrained industrial tables, LLMs are more effective as complementary components than as replacements for strong tabular baselines.