paper-with-me

Papers

Quantifying and Reducing Speaker Heterogeneity within the Common Voice Corpus for Phonetic Analysis

2025-05-31 · Miao Zhang, Aref Farhadipour, Annie Baker, Jiachen Ma, Bogdan Pricop, Eleanor Chodroff

With its crosslinguistic and cross-speaker diversity, the Mozilla Common Voice Corpus (CV) has been a valuable resource for multilingual speech technology and holds tremendous potential for research in crosslinguistic phonetics and speech sciences. Properly accounting for speaker variation is, however, key to the theoretical and statistical bases of speech research. While CV provides a client ID as an approximation to a speaker ID, multiple speakers can contribute under the same ID. This study aims to quantify and reduce heterogeneity in the client ID for a better approximation of a true, though still anonymous speaker ID. Using ResNet-based voice embeddings, we obtained a similarity score among recordings with the same client ID, then implemented a speaker discrimination task to identify an optimal threshold for reducing perceived speaker heterogeneity. These results have major downstream applications for phonetic analysis and the development of speaker-based speech technology.

📄 PDF Abstract BibTeX arXiv:2506.00733

Code (0)

등록된 구현이 없습니다.

Tasks

Diversity

Similar Papers 제목 키워드 기반

Measuring Heterogeneity in Machine Learning with Distributed Energy Distance

2025-01-27 · Mengchen Fan, Baocheng Geng, Roman Shterenberg, Joseph A. Casey 외

In distributed and federated learning, heterogeneity across data sources remains a major obstacle to effective model aggregation and convergence. We focus on feature heterogeneity and introduce energy distance as a sensi…

Federated Learning

Libri-Adapt: A New Speech Dataset for Unsupervised Domain Adaptation

2020-09-06 · Akhil Mathur, Fahim Kawsar, Nadia Berthouze, Nicholas D. Lane

This paper introduces a new dataset, Libri-Adapt, to support unsupervised domain adaptation research on speech recognition models. Built on top of the LibriSpeech corpus, Libri-Adapt contains English speech recorded on m…

Domain Adaptationspeech-recognitionSpeech RecognitionUnsupervised Domain Adaptation

Triplet loss based embeddings for forensic speaker identification in Spanish

2021-02-24 · Emmanuel Maqueda, Javier Alvarez-Jimenez, Carlos Mena, Ivan Meza

With the advent of digital technology, it is more common that committed crimes or legal disputes involve some form of speech recording where the identity of a speaker is questioned [1]. In face of this situation, the fie…

Speaker IdentificationTriplet

CASL-VAE: Learning Structured Latent Variables from Unpaired Data for Semi-supervised Clustering and Paired Sample Generation

2026-07-09 · Sai Spandana Chintapalli, Pratik Chaudhari, Christos Davatzikos arxiv

Quantifying variability in a target population relative to a reference population is central to many scientific and clinical problems (e.g., diseased vs. healthy). Yet, without paired data and in the presence of heteroge…

Commonality and Individuality! Integrating Humor Commonality with Speaker Individuality for Humor Recognition

2025-02-07 · Haohao Zhu, Junyu Lu, Zeyuan Zeng, Zewen Bai 외

Humor recognition aims to identify whether a specific speaker's text is humorous. Current methods for humor recognition mainly suffer from two limitations: (1) they solely focus on one aspect of humor commonalities, igno…