paper-with-me

Papers

Rhythm Features for Speaker Identification

2025-06-07 · Nick Mehlman, Thomas Thebaud, Dani Byrd, Shri Narayanan

While deep learning models have demonstrated robust performance in speaker recognition tasks, they primarily rely on low-level audio features learned empirically from spectrograms or raw waveforms. However, prior work has indicated that idiosyncratic speaking styles heavily influence the temporal structure of linguistic units in speech signals (rhythm). This makes rhythm a strong yet largely overlooked candidate for a speech identity feature. In this paper, we test this hypothesis by applying deep learning methods to perform text-independent speaker identification from rhythm features. Our findings support the usefulness of rhythmic information for speaker recognition tasks but also suggest that high intra-subject variability in ad-hoc speech can degrade its effectiveness.

📄 PDF Abstract BibTeX arXiv:2506.06834

Code (0)

등록된 구현이 없습니다.

Tasks

Deep LearningRhythmSpeaker IdentificationSpeaker Recognition

Similar Papers 제목 키워드 기반

Speech Rhythm-Based Speaker Embeddings Extraction from Phonemes and Phoneme Duration for Multi-Speaker Speech Synthesis

2024-02-11 · Kenichi Fujita, Atsushi Ando, Yusuke Ijima

This paper proposes a speech rhythm-based method for speaker embeddings to model phoneme duration using a few utterances by the target speaker. Speech rhythm is one of the essential factors among speaker characteristics,…

RhythmSpeaker IdentificationSpeech Synthesis

Arabic Speech Rhythm Corpus: Read and Spontaneous Speaking Styles

2020-05-01 · LREC 2020 5 · Omnia Ibrahim, Homa Asadi, Eman Kassem, Volker Dellwo

Databases for studying speech rhythm and tempo exist for numerous languages. The present corpus was built to allow comparisons between Arabic speech rhythm and other languages. 10 Egyptian speakers (gender-balanced) prod…

RhythmSpeaker Recognition

Zero-shot text-to-speech synthesis conditioned using self-supervised speech representation model

2023-04-24 · Kenichi Fujita, Takanori Ashihara, Hiroki Kanagawa, Takafumi Moriya 외

This paper proposes a zero-shot text-to-speech (TTS) conditioned by a self-supervised speech-representation model acquired through self-supervised learning (SSL). Conventional methods with embedding vectors from x-vector…

RhythmSelf-Supervised LearningSpeech Synthesistext-to-speech+2

Analyzing and Improving Speaker Similarity Assessment for Speech Synthesis

2025-07-02 · Marc-André Carbonneau, Benjamin van Niekerk, Hugo Seuté, Jean-Philippe Letendre 외 arxiv

Modeling voice identity is challenging due to its multifaceted nature. In generative speech systems, identity is often assessed using automatic speaker verification (ASV) embeddings, designed for discrimination rather th…

Speaker VerificationSpeech Synthesis

AS-Speech: Adaptive Style For Speech Synthesis

2024-09-09 · Zhipeng Li, Xiaofen Xing, Jun Wang, Shuaiqi Chen 외

In recent years, there has been significant progress in Text-to-Speech (TTS) synthesis technology, enabling the high-quality synthesis of voices in common scenarios. In unseen situations, adaptive TTS requires a strong g…

RhythmSpeech Synthesistext-to-speechText to Speech+1