慣性聚合 高效追蹤和閱讀你感興趣的部落格、新聞、科技資訊
閱讀原文 在慣性聚合中打開

推薦訂閱源

N
Netflix TechBlog - Medium
Microsoft Azure Blog
Microsoft Azure Blog
freeCodeCamp Programming Tutorials: Python, JavaScript, Git & More
S
Security Archives - TechRepublic
Cyberwarzone
Cyberwarzone
D
Darknet – Hacking Tools, Hacker News & Cyber Security
IntelliJ IDEA : IntelliJ IDEA – the Leading IDE for Professional Development in Java and Kotlin | The JetBrains Blog
IntelliJ IDEA : IntelliJ IDEA – the Leading IDE for Professional Development in Java and Kotlin | The JetBrains Blog
博客园 - 【当耐特】
A
About on SuperTechFans
T
ThreatConnect
IT之家
IT之家
阮一峰的网络日志
阮一峰的网络日志
B
Blog
T
Tailwind CSS Blog
G
GRAHAM CLULEY
F
Future of Privacy Forum
V
Vulnerabilities – Threatpost
J
Java Code Geeks
量子位
博客园 - 叶小钗
Last Week in AI
Last Week in AI
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻
Stack Overflow Blog
Stack Overflow Blog
李成银的技术随笔
Spread Privacy
Spread Privacy
The Hacker News
The Hacker News
S
Schneier on Security
T
True Tiger Recordings
Vercel News
Vercel News
C
CXSECURITY Database RSS Feed - CXSecurity.com
C
Cybersecurity and Infrastructure Security Agency CISA
Latest news
Latest news
F
Fox-IT International blog
The Register - Security
The Register - Security
MongoDB | Blog
MongoDB | Blog
博客园 - 聂微东
Threat Intelligence Blog | Flashpoint
Threat Intelligence Blog | Flashpoint
Know Your Adversary
Know Your Adversary
GbyAI
GbyAI
L
LangChain Blog
MyScale Blog
MyScale Blog
AWS News Blog
AWS News Blog
D
Docker
小众软件
小众软件
Stack Overflow Blog
Stack Overflow Blog
Microsoft Security Blog
Microsoft Security Blog
T
Tor Project blog
T
The Exploit Database - CXSecurity.com
P
Palo Alto Networks Blog
Malwarebytes
Malwarebytes

cs.LG updates on arXiv.org

Personalized Generative Models for Contextual Debiasing From Privacy to Generalization: Linear Max-Information Bounds for DP-SGD When Does Deep RL Beat Calibrated Baselines? A Benchmark Study on Adaptive Resource Control Amortized Factor Inference Networks for Posterior Inference Classification and detection of multiple UAVs using rational Gaussian wavelet neural networks Planning Neural Dynamics with Lie Group Embedding through Supervised Projective Manifold Learning Modeling Dynamic Mixtures of Time-Delay Systems from Streaming Time Series AirCast-SR: A Foundation Model for Kilometer-Scale Atmospheric Super-Resolution via Latent Consistency Diffusion Neural Bayesian Sequential Routing GAC: Noise-Aware Adaptive Mixing for Hybrid SFT-RL Post-Training Provably Communication-Efficient and Privacy-Preserving Federated Graph Neural Networks Function-Valued Causal Influence in Nonlinear Time Series The Bridge-Garden Dilemma in LLM Distillation: Why Mixing Hard and Soft Labels Works Balancing Plasticity and Stability with Fast and Slow Successor Features InfoQuant: Shaping Activation Distributions for Low-Bit LLM Quantization FM-fMRI: Event Conditioned Flow Matching for Rest-to-Task fMRI Time-Series Synthesis TrackRef3D: Multi-View Consistent Track-then-Label for Open-World Referring Segmentation in 3D Gaussian Splatting TSFMAudit: Data Contamination Auditing in Forecasting Time Series Foundation Models On the Push-Based Asynchronous Federated Learning: A Bias-Correction Aggregation Approach CSV-ViT: A Vision Transformer with the Variable-sized Cortical Supervertices for Detection of Alzheimer's Disease Pathologies Max-Window Scale Estimation for Near-Lossless HiF8 W8A8 Quantization-Aware Training Online Learning on Hidden-Convex Losses via Algorithmic Equivalence: Optimal Regret, Geometric Barrier, and Bandit Feedback Curriculum Learning for Safety Alignment A PAC-Bayesian View of Generalisation for Physics-Informed Machine Learning Dynamic Link Prediction with Temporally Enhanced Signed Graph Neural Networks GEM: Geometric Entropy Mixing for Optimal LLM Data Curation Reparametrizing Shampoo and SOAP for Subspace Basis Updates and BFloat16 Storage Unified Neural Scaling Laws Semigroup Consistency as a Diagnostic for Learned Physics Simulators QAM-W: Joint 2D Codebook Quantization for LLM Weights via Hadamard Rotation and Activation-Aware Scaling HRVConformer: Neonatal Hypoxic-Ischemic Encephalopathy Classification from the Heart Rate signals Scaling World-Model Reinforcement Learning Through Diffusion Policy Optimization A Hybrid Vision-Language Architecture for Automated Defect Reasoning and Report Generation in Industrial Inspection Quantized Keys Steal Attention: Bias Correction for KV-Cache Compression in Video Diffusion BioFact-MoE: Biologically Factorized Mixture of Experts for Vision-Language Prognostic Modeling in Hepatocellular Carcinoma Bridging Classification and Reconstruction: Cooperative Time Series Anomaly Detection SilIF: Silhouette-Augmented Isolation Forest for Unsupervised Transaction Fraud Detection Co-folding model guided by structural proteomics Energy-Gated Attention and Wavelet Positional Encoding: Complementary Inductive Biases for Transformer Attention Stateful Inference for Low-Latency Multi-Agent Tool Calling Two-Parameter Flows for Learning Population Dynamics of Physical Systems On the Role of Inductive Bias in Time-Series Pretraining: A Case Study in Learning Generalizable Representations for Clinical Time Series A Fast and Generic Energy-Shifting Transformer for Hybrid Monte Carlo Radiotherapy Calculation ARBITER: Reasoning Trajectory Basins and Majority Vote Failures in Test-Time Sampling When Rule Violations Are Rare: Chimera Training for Logical Anomaly Detection The Constraint Tax: Measuring Validity-Correctness Tradeoffs in Structured Outputs for Small Language Models When Correct Demonstrations Hurt: Rethinking the Role of Exemplars in In-Context Learning Benchmarking Convolutional, Transformer, Hybrid, and Vision Language Models for Multi Disease Retinal Screening Rotation-Invariant Spherical Watermarking via Third-Order SO(3) Representation Coupling MULTISEISMO: A Multimodal Seismic Dataset and Model for Cross-Modal Seismic Understanding
密集检索的測試時計算:使用凍結嵌入模型的代理程式生成
Han Xiao · 2026-05-13 · via cs.LG updates on arXiv.org

檢視 PDF HTML (實驗性)

摘要:在測試時進行計算普遍被認為只對大型推理模型有益。我們證明它也對小型嵌入模型有幫助。由於現代嵌入模型是由大語言模型主幹衍生出來的,一個凍結的編碼器應該能夠從額外的推理計算中受益而不需要重新訓練。一個代理程序搜索迴圈在凍結的編碼器 API 上探索了 144 候選程序,並生產了十二個帕累托優化程序,涵蓋了從 $c=1.2$ 到 $14.7$ 的成本比率,超過單次通過基線。搜索獨立地重新發現了 Rocchio 虛假相關反饋、ColBERT 風格的句子粒度 MaxSim、互惠排名融合,以及 Fisher 線性判別,所有這些都不需要可訓練參數或外部模型。每一個邊緣程序在所有 14 個 MMTEB 检索任務上,涵蓋法律、金融、長文檔和一般領域,都提高了 nDCG@10,超過凍結的基線。這些程序無需修改即可轉移到未見過的編碼器家族和十九個保留的检索任務,68% 的模型-任務對承認至少有一個邊緣程序,其性能優於餘弦基線。
評論: 16 頁,4 個圖
主題: 機器學習 (cs.LG); 計算與語言 (cs.CL); 資訊搜尋 (cs.IR)
引用格式: arXiv:2605.11374 [cs.LG]
  (或 arXiv:2605.11374v3 [cs.LG]) for this version)
  https://doi.org/10.48550/arXiv.2605.11374

arXiv發行的DOI透過DataCite

提交通訊錄

從: Han Xiao [查看郵件]
[v1] 周二,2026年5月12日 00:56:34 UTC (215 KB)
[v2] Wed, 13 May 2026 00:56:03 UTC (126 KB)
[v3] Tue, 26 May 2026 14:57:14 UTC (264 KB)