惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

美团技术团队
J
Java Code Geeks
有赞技术团队
有赞技术团队
GbyAI
GbyAI
奇客Solidot–传递最新科技情报
奇客Solidot–传递最新科技情报
酷 壳 – CoolShell
酷 壳 – CoolShell
博客园 - 叶小钗
阮一峰的网络日志
阮一峰的网络日志
Microsoft Security Blog
Microsoft Security Blog
IT之家
IT之家
G
Google Developers Blog
月光博客
月光博客
Cyber Security Advisories - MS-ISAC
Cyber Security Advisories - MS-ISAC
S
SegmentFault 最新的问题
博客园 - 三生石上(FineUI控件)
钛媒体:引领未来商业与生活新知
钛媒体:引领未来商业与生活新知
博客园 - Franky
腾讯CDC
V
Visual Studio Blog
博客园 - 【当耐特】
D
Docker
H
Hackread – Cybersecurity News, Data Breaches, AI and More
Engineering at Meta
Engineering at Meta
L
LangChain Blog

eess.AS updates on arXiv.org

Dependence on Early and Late Reverberation of Single-Channel Speaker Distance Estimation MIST: Multimodal Interactive Speech-based Tool-calling Conversational Assistants for Smart Homes LiVeAction: a Lightweight, Versatile, and Asymmetric Neural Codec Design for Real-time Operation Weight-Decay Turns Transformer Loss Landscapes Villani: Functional-Analytic Foundations for Optimization and Generalization PairAlign: A Framework for Sequence Tokenization via Self-Alignment with Applications to Audio Tokenization WavCube: Unifying Speech Representation for Understanding and Generation via Semantic-Acoustic Joint Modeling Predictive-Generative Drift Decomposition for Speech Enhancement and Separation Minimizing Modality Gap from the Input Side: Your Speech LLM Can Be a Prosody-Aware Text LLM X-Voice: Enabling Everyone to Speak 30 Languages via Zero-Shot Cross-Lingual Voice Cloning JASTIN: Aligning LLMs for Zero-Shot Audio and Speech Evaluation via Natural Language Instructions Phoneme-Level Deepfake Detection Across Emotional Conditions Using Self-Supervised Embeddings When Audio-Language Models Fail to Leverage Multimodal Context for Dysarthric Speech Recognition Dimensionality-Aware Anomaly Detection in Learned Representations of Self-Supervised Speech Models Mitigating Multimodal LLMs Hallucinations via Relevance Propagation at Inference Time Virtual Speech Therapist: A Clinician-in-the-Loop AI Speech Therapy Agent for Personalized and Supervised Therapy LASE: Language-Adversarial Speaker Encoding for Indic Cross-Script Identity Preservation Towards Improving Speaker Distance Estimation through Generative Impulse Response Augmentation Beyond Decodability: Reconstructing Language Model Representations with an Encoding Probe MMAudioReverbs: Video-Guided Acoustic Modeling for Dereverberation and Room Impulse Response Estimation Alethia: A Foundational Encoder for Voice Deepfakes From Birdsong to Rumbles: Classifying Elephant Calls with Out-of-Species Embeddings Beyond the Baseband: Adaptive Multi-Band Encoding for Full-Spectrum Bioacoustics Classification Predicting Upcoming Stuttering Events from Three-Second Audio: Stratified Evaluation Reveals Severity-Selective Precursors, and the Model Deploys Fully On-Device The False Resonance: A Critical Examination of Emotion Embedding Similarity for Speech Generation Evaluation DiffAnon: Diffusion-based Prosody Control for Voice Anonymization Recurrence-Based Nonlinear Vocal Dynamics as Digital Biomarkers for Depression Detection from Conversational Speech One Voice, Many Tongues: Cross-Lingual Voice Cloning for Scientific Speech Similarity Choice and Negative Scaling in Supervised Contrastive Learning for Deepfake Audio Detection Walking Through Uncertainty: An Empirical Study of Uncertainty Estimation for Audio-Aware Large Language Models Praxy Voice: Voice-Prompt Recovery + BUPS for Commercial-Class Indic TTS from a Frozen Non-Indic Base at Zero Commercial-Training-Data Cost
Subspace Track-before-Detect for Passive Multi-Target Tra...
[Submitted on 25 May 2026 (v1), last revised 1 Jul 2026 (this ve · 2026-05-25 · via eess.AS updates on arXiv.org

View PDF HTML (experimental)

Abstract:Passive multi-target tracking (MTT) aims to infer the time-varying kinematic and activity states of an unknown number of sources that emit unknown and possibly nonstationary signals, using only noisy mixtures of these signals observed at sensors. Track-before-detect (TBD) methods improve noise robustness by evaluating multi-target hypotheses directly on raw sensor data, without relying on a preceding detection stage. However, existing TBD likelihoods typically assume that the contribution of each active target to the observation is determined solely by its kinematic state. This assumption does not hold in passive sensing scenarios, where the observed mixtures also depend on unknown and possibly nonstationary source signals.
To address this issue, we propose subspace TBD, a passive multi-target TBD method that employs a source-signal-insensitive likelihood derived from the complex spherical Student's $t$ (cST) distribution. Instead of explicitly modeling or estimating the nuisance source signals, the method represents each multi-target hypothesis by the subspace spanned by source steering vectors. The cST likelihood then evaluates how well the normalized multichannel mixtures align with this subspace. We conducted acoustic MTT simulations with two moving speakers in noisy, reverberant environments, comparing the proposed method with a baseline consisting of steered response power with phase transform (SRP-PHAT) followed by a sequential Monte Carlo implementation of the generalized labeled multi-Bernoulli filter (SMC-GLMB). The proposed method achieved lower mean optimal subpattern assignment (OSPA) values in all tested conditions.

Submission history

From: Nobutaka Ito B.E. M.E. Ph.D. [view email]
[v1] Mon, 25 May 2026 06:57:38 UTC (999 KB)
[v2] Wed, 1 Jul 2026 13:00:56 UTC (1,743 KB)