paper-with-me

Papers

Variable frame rate-based data augmentation to handle speaking-style variability for automatic speaker verification

2020-08-08 · Amber Afshan, Jinxi Guo, Soo Jin Park, Vijay Ravi, Alan McCree, Abeer Alwan

The effects of speaking-style variability on automatic speaker verification were investigated using the UCLA Speaker Variability database which comprises multiple speaking styles per speaker. An x-vector/PLDA (probabilistic linear discriminant analysis) system was trained with the SRE and Switchboard databases with standard augmentation techniques and evaluated with utterances from the UCLA database. The equal error rate (EER) was low when enrollment and test utterances were of the same style (e.g., 0.98% and 0.57% for read and conversational speech, respectively), but it increased substantially when styles were mismatched between enrollment and test utterances. For instance, when enrolled with conversation utterances, the EER increased to 3.03%, 2.96% and 22.12% when tested on read, narrative, and pet-directed speech, respectively. To reduce the effect of style mismatch, we propose an entropy-based variable frame rate technique to artificially generate style-normalized representations for PLDA adaptation. The proposed system significantly improved performance. In the aforementioned conditions, the EERs improved to 2.69% (conversation -- read), 2.27% (conversation -- narrative), and 18.75% (pet-directed -- read). Overall, the proposed technique performed comparably to multi-style PLDA adaptation without the need for training data in different speaking styles per speaker.

📄 PDF Abstract BibTeX arXiv:2008.03616

Code (0)

등록된 구현이 없습니다.

Tasks

Data AugmentationSpeaker Verification

Similar Papers 제목 키워드 기반

Practical Wi-Fi-based Motion Recognition Under Variable Traffic Patterns

2026-05-08 · Guolin Yin, Junqing Zhang, Guanxiong Shen, Simon L. Cotton arxiv

Wi-Fi sensing detects human motions and activities by analysing the channel state information (CSI) derived from Wi-Fi transmissions. However, the impact of variable transmission traffic, which dictates the effective sam…

Activity Recognition

Learning to Retrieve for Environmental Knowledge Discovery: An Augmentation-Adaptive Self-Supervised Learning Framework

2025-09-18 · Shiyuan Luo, Runlong Yu, Chonghao Qiu, Rahul Ghosh 외 arxiv

The discovery of environmental knowledge depends on labeled task-specific data, but is often constrained by the high cost of data collection. Existing machine learning approaches usually struggle to generalize in data-sp…

Self-Supervised LearningData Augmentation

Exact Regular-Constrained Variable-Order Markov Generation via Sparse Context-State Belief Propagation

2026-05-08 · François Pachet arxiv

Variable-order Markov models generate sequences over a finite alphabet by conditioning each symbol on the longest available suffix of the generated history. Regular constraints, by contrast, describe finite-horizon contr…

Data Augmentation

Bayesian Non-Homogeneous Markov Models via Polya-Gamma Data Augmentation with Applications to Rainfall Modeling

2017-01-11 · Tracy Holsclaw, Arthur M. Greene, Andrew W. Robertson, Padhraic Smyth

Discrete-time hidden Markov models are a broadly useful class of latent-variable models with applications in areas such as speech recognition, bioinformatics, and climate data analysis. It is common in practice to introd…

Data Augmentationspeech-recognitionSpeech Recognition

Variable-Speed Teaching-Playback as Real-World Data Augmentation for Imitation Learning

2024-12-04 · Nozomu Masuya, Hiroshi Sato, Koki Yamane, Takuya Kusume 외

Because imitation learning relies on human demonstrations in hard-to-simulate settings, the inclusion of force control in this method has resulted in a shortage of training data, even with a simple change in speed. Altho…

Data AugmentationImitation LearningPositionRobot Manipulation