惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

量子位
aimingoo的专栏
aimingoo的专栏
C
CXSECURITY Database RSS Feed - CXSecurity.com
Stack Overflow Blog
Stack Overflow Blog
C
CERT Recently Published Vulnerability Notes
T
Tailwind CSS Blog
腾讯CDC
罗磊的独立博客
Security Latest
Security Latest
K
Kaspersky official blog
A
Arctic Wolf
博客园 - Franky
D
Docker
博客园 - 司徒正美
GbyAI
GbyAI
T
Tenable Blog
Engineering at Meta
Engineering at Meta
A
About on SuperTechFans
H
Help Net Security
Threat Intelligence Blog | Flashpoint
Threat Intelligence Blog | Flashpoint
L
Lohrmann on Cybersecurity
小众软件
小众软件
V
V2EX
T
Threatpost
T
Threat Research - Cisco Blogs
T
The Exploit Database - CXSecurity.com
P
Palo Alto Networks Blog
P
Privacy & Cybersecurity Law Blog
S
Securelist
Google DeepMind News
Google DeepMind News
I
Intezer
The Register - Security
The Register - Security
NISL@THU
NISL@THU
L
LINUX DO - 热门话题
C
Cisco Blogs
AWS News Blog
AWS News Blog
MyScale Blog
MyScale Blog
S
Schneier on Security
Scott Helme
Scott Helme
T
The Blog of Author Tim Ferriss
G
Google Developers Blog
Project Zero
Project Zero
Cyberwarzone
Cyberwarzone
CTFtime.org: upcoming CTF events
CTFtime.org: upcoming CTF events
钛媒体:引领未来商业与生活新知
钛媒体:引领未来商业与生活新知
I
InfoQ
Cisco Talos Blog
Cisco Talos Blog
Know Your Adversary
Know Your Adversary
L
LangChain Blog
P
Proofpoint News Feed

Apple Machine Learning Research

Environment-free Synthetic Data Generation for API-Calling Agents Accelerating Text-to-Video Generation with Calibrated Sparse Attention RayRoPE: Projective Ray Positional Encoding for Multi-View Attention LVSum: A Benchmark for Timestamp-Aware Long Video Summarization Length Value Model: Scalable Value Pretraining for Token-Level Length Modeling When Unlearning Is Free: Leveraging Low Influence Points to Reduce Computational Costs Show Me Examples: Inferring Visual Concepts from Image Sets Location-Invariant Properties of Functions Versus Properties of Distributions: United in Testing but Separated in Verification Interactive Proofs for General Distribution Properties Doubly Sub-linear Interactive Proofs of Proximity Personalizing Incremental Video Search with Hybrid Text and ID Embeddings Embarrassingly Simple Self-Distillation Improves Code Generation CLaRa: Bridging Retrieval and Generation with Continuous Latent Reasoning Uncertainty Quantification for LLM Function-Calling One Layer Is Enough: Adapting Pretrained Visual Encoders for Image Generation Proactive Agent Research Environment: Simulating Active Users to Evaluate Proactive Assistants Multilingual Semantic Retrieval for Apple Music Search Behavioral Privacy Leakage in Agentic Negotiation: Formalizing and Mitigating Inference Attacks via Randomized Policies Incentivizing Temporal-Awareness in Egocentric Video Understanding Models Recursive Language Models Meet Uncertainty: The Surprising Effectiveness of Self-Reflective Program Search for Long Context Unmasking On-Policy Distillation: Where It Helps, Where It Hurts, and Why Taming Text-to-Sounding Video Generation via Advanced Modality Condition and Interaction DynaMiCS: Fine-Tuning LLMs with Performance Constraints Using Dynamic Mixtures LensVLM: Selective Context Expansion for Compressed Visual Representation of Text MT-EditFlow: Reinforcement Learning for Multi-Turn Image Editing with Flow Matching Weblica: Scalable and Reproducible Training Environments for Visual Web Agents FlowEval: Reference-Based Evaluation of Generated User Interfaces A Single Neuron Is Sufficient to Bypass Safety Alignment in Large Language Models Scaling Properties of Continuous Diffusion Spoken Language Models Path-Constrained Mixture-of-Experts TopoPrimer: The Missing Topological Context in Forecasting Models Multi-Agent Teams Hold Experts Back VideoFlexTok: Flexible-Length Coarse-to-Fine Video Tokenization Amortizing Maximum Inner Product Search with Learned Support Functions On Robustness and Chain-of-Thought Consistency of RL-Finetuned VLMs MemoryLLM: Plug-n-Play Interpretable Feed-Forward Memory for Transformers Learning Structured Reasoning via Tractable Trajectory Control Learning Unmasking Policies for Diffusion Language Models Residual Context Diffusion Language Models Conformal Thinking: Risk Control for Reasoning on a Compute Budget Anti-Causal Domain Generalization: Leveraging Unlabeled Data Metric-Dependent Annotation Saturation for Learning from Label Distributions Nine Judges, Two Effective Votes: Correlated Errors Undermine LLM Evaluation Panels Introducing the Third Generation of Apple’s Foundation Models IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) 2026 VSAS-Bench: Real-Time Evaluation of Visual Streaming Assistant Models EpiCache: Episodic KV Cache Management for Long-Term Conversation on Resource-Constrained Environments BalCapRL: A Balanced Framework for RL-Based MLLM Image Captioning Apple Workshop on Privacy-Preserving Machine Learning & AI 2026 Velox: Learning Representations of 4D Geometry and Appearance RVPO: Risk-Sensitive Alignment via Variance Regularization Large-Scale High-Quality 3D Gaussian Head Reconstruction from Multi-View Captures Text-Conditional JEPA for Learning Semantically Rich Visual Representations What Matters in Practical Learned Image Compression SpecMD: A Comprehensive Study on Speculative Expert Prefetching From Where Things Are to What They’re For: Benchmarking Spatial–Functional Intelligence for Multimodal LLMs STARFlow-V: End-to-End Video Generative Modeling with Normalizing Flows Bootstrapping Sign Language Annotations with Sign Language Models International Conference on Acoustics, Speech and Signal Processing (ICASSP) 2026 Adaptive Thinking: Large Language Models Know When to Think in Latent Space DSO: Direct Steering Optimization for Bias Mitigation StereoFoley: Object-Aware Stereo Audio Generation from Video LaDiR: Latent Diffusion Enhances LLMs for Text Reasoning Local Mechanisms of Compositional Generalization in Conditional Diffusion Learning Long-Term Motion Embeddings for Efficient Kinematics Generation ParaRNN: Large-Scale Nonlinear RNNs, Trainable in Parallel Apple Machine Learning Research at ICLR 2026 Can Large Language Models Understand Context? International Conference on Learning Representations (ICLR) 2026 Cram Less to Fit More: Training Data Pruning Improves Memorization of Facts Efficient Privacy Loss Accounting for Subsampling and Random Allocation ACM Human-Computer Interaction Conference (CHI) 2026 A Theoretical Framework for Acoustic Neighbor Embeddings Governance-Aware Agent Telemetry for Closed-Loop Enforcement in Multi-Agent AI Systems SQUIRE: Interactive UI Authoring via Slot QUery Intermediate REpresentations Personalized Group Relative Policy Optimization for Heterogenous Preference Alignment ProText: A Benchmark Dataset for Measuring (Mis)gendering in Long-Form Texts Beyond Real Data: Synthetic Data through the Lens of Regularization Entropy-Preserving Reinforcement Learning Less Gaussians, Texture More: 4K Feed-Forward Textured Splatting
Revisiting ASR Error Correction with Specialized Models
2026-07-06 · via Apple Machine Learning Research

AuthorsZijin Gu, Tatiana Likhomanenko, Richard He Bai, Erik McDermott, Ronan Collobert, Navdeep Jaitly†**

Language models play a central role in automatic speech recognition (ASR), yet most methods rely on text-only models unaware of ASR error patterns. Recently, large language models (LLMs) have been applied to ASR correction, but introduce latency and hallucination concerns. We revisit ASR error correction with compact seq2seq models, trained on ASR errors from real and synthetic audio. To scale training, we construct synthetic corpora via cascaded TTS and ASR, finding that matching the diversity of realistic error distributions is key. We propose correction-first decoding, where the correction model generates candidates rescored using ASR acoustic scores. With 15x fewer parameters than LLMs, our model achieves 1.5/3.3% WER on LibriSpeech test-clean/other, outperforms LLMs, generalizes across ASR architectures (CTC, Seq2seq, Transducer) and diverse domains, and provides precise corrections in the low-error regime where LLMs struggle.

  • † Google
  • ** Work done while at Apple

Related readings and updates.

This paper presents an efficient decoding approach for end-to-end automatic speech recognition (E2E-ASR) with large language models (LLMs). Although shallow fusion is the most common approach to incorporate language models into E2E-ASR decoding, we face two practical problems with LLMs. (1) LLM inference is computationally costly. (2) There may be a vocabulary mismatch between the ASR model and the LLM. To resolve this mismatch, we need to…

Read more

In recent years, end-to-end automatic speech recognition (ASR) systems have proven themselves remarkably accurate and performant, but these systems still have a significant error rate for entity names which appear infrequently in their training data. In parallel to the rise of end-to-end ASR systems, large language models (LLMs) have proven to be a versatile tool for various natural language processing (NLP) tasks. In NLP tasks where a database…

Read more