paper-with-me

Papers

Analysis of Speech Temporal Dynamics in the Context of Speaker Verification and Voice Anonymization

2024-12-22 · Natalia Tomashenko, Emmanuel Vincent, Marc Tommasi

In this paper, we investigate the impact of speech temporal dynamics in application to automatic speaker verification and speaker voice anonymization tasks. We propose several metrics to perform automatic speaker verification based only on phoneme durations. Experimental results demonstrate that phoneme durations leak some speaker information and can reveal speaker identity from both original and anonymized speech. Thus, this work emphasizes the importance of taking into account the speaker's speech rate and, more importantly, the speaker's phonetic duration characteristics, as well as the need to modify them in order to develop anonymization systems with strong privacy protection capacity.

📄 PDF Abstract BibTeX arXiv:2412.17164

Code (0)

등록된 구현이 없습니다.

Tasks

Speaker Verification

Similar Papers 제목 키워드 기반

Exploiting Context-dependent Duration Features for Voice Anonymization Attack Systems

2025-07-21 · Natalia Tomashenko, Emmanuel Vincent, Marc Tommasi arxiv

The temporal dynamics of speech, encompassing variations in rhythm, intonation, and speaking rate, contain important and unique information about speaker identity. This paper proposes a new method for representing speake…

Speaker Verification

Speaker-Reasoner: Scaling Interaction Turns and Reasoning Patterns for Timestamped Speaker-Attributed ASR

2026-04-03 · Zhennan Lin, Shuai Wang, Zhaokai Sun, Pengyuan Xie 외 arxiv

Transcribing and understanding multi-speaker conversations requires speech recognition, speaker attribution, and timestamp localization. While speech LLMs excel at single-speaker tasks, multi-speaker scenarios remain cha…

Speech Recognition

DNN driven Speaker Independent Audio-Visual Mask Estimation for Speech Separation

2018-07-31 · Mandar Gogate, Ahsan Adeel, Ricard Marxer, Jon Barker 외

Human auditory cortex excels at selectively suppressing background noise to focus on a target speaker. The process of selective attention in the brain is known to contextually exploit the available audio and visual cues …

Speech Separation

Optimization of Speaker Extraction Neural Network with Magnitude and Temporal Spectrum Approximation Loss

2019-03-24 · Cheng-Lin Xu, Wei Rao, Eng Siong Chng, Haizhou Li

The SpeakerBeam-FE (SBF) method is proposed for speaker extraction. It attempts to overcome the problem of unknown number of speakers in an audio recording during source separation. The mask approximation loss of SBF is …

KM-Speaker: Keypoint-Based Style Control for High-Quality Speech-Driven 3D Facial Animation and Dialogue Localization

2026-06-26 · Arthur Josi, Emeline Got, Abdallah Dib, Luiz Gustavo Hafemann 외 arxiv

Speech-driven 3D facial animation methods face significant challenges in simultaneously achieving high-fidelity motion and precise artistic control at production quality. Existing controllable models typically learn glob…