paper-with-me

Papers

Speaker Embeddings to Improve Tracking of Intermittent and Moving Speakers

2025-06-23 · Taous Iatariene, Can Cui, Alexandre Guérin, Romain Serizel

Speaker tracking methods often rely on spatial observations to assign coherent track identities over time. This raises limits in scenarios with intermittent and moving speakers, i.e., speakers that may change position when they are inactive, thus leading to discontinuous spatial trajectories. This paper proposes to investigate the use of speaker embeddings, in a simple solution to this issue. We propose to perform identity reassignment post-tracking, using speaker embeddings. We leverage trajectory-related information provided by an initial tracking step and multichannel audio signal. Beamforming is used to enhance the signal towards the speakers' positions in order to compute speaker embeddings. These are then used to assign new track identities based on an enrollment pool. We evaluate the performance of the proposed speaker embedding-based identity reassignment method on a dataset where speakers change position during inactivity periods. Results show that it consistently improves the identity assignment performance of neural and standard tracking systems. In particular, we study the impact of beamforming and input duration for embedding extraction.

📄 PDF Abstract BibTeX arXiv:2506.19875

Code (0)

등록된 구현이 없습니다.

Tasks

Position

Similar Papers 제목 키워드 기반

Tracking of Intermittent and Moving Speakers : Dataset and Metrics

2025-06-11 · Taous Iatariene, Alexandre Guérin, Romain Serizel

This paper presents the problem of tracking intermittent and moving sources, i.e, sources that may change position when they are inactive. This issue is seldom explored, and most current tracking methods rely on spatial …

Position

Online Binaural Speech Separation of Moving Speakers With a Wavesplit Network

2023-03-13 · Cong Han, Nima Mesgarani

Binaural speech separation in real-world scenarios often involves moving speakers. Most current speech separation methods use utterance-level permutation invariant training (u-PIT) for training. In inference time, howeve…

Online ClusteringSpeaker SeparationSpeech Separation

ImagineNET: Target Speaker Extraction with Intermittent Visual Cue through Embedding Inpainting

2022-10-31 · Zexu Pan, Wupeng Wang, Marvin Borsdorf, Haizhou Li

The speaker extraction technique seeks to single out the voice of a target speaker from the interfering voices in a speech mixture. Typically an auxiliary reference of the target speaker is used to form voluntary attenti…

Target Speaker Extraction

Adaptive Rotary Steering with Joint Autoregression for Robust Extraction of Closely Moving Speakers in Dynamic Scenarios

2026-01-18 · Jakob Kienegger, Timo Gerkmann arxiv

Latest advances in deep spatial filtering for Ambisonics demonstrate strong performance in stationary multi-speaker scenarios by rotating the sound field toward a target speaker prior to multi-channel enhancement. For ap…

Multi-Modal Perception Attention Network with Self-Supervised Learning for Audio-Visual Speaker Tracking

2021-12-14 · Yidi Li, Hong Liu, Hao Tang

Multi-modal fusion is proven to be an effective method to improve the accuracy and robustness of speaker tracking, especially in complex scenarios. However, how to combine the heterogeneous information and exploit the co…

Self-Supervised Learning