paper-with-me

Papers

Window Size Versus Accuracy Experiments in Voice Activity Detectors

2026-01-24 · Max McKinnon, Samir Khaki, Chandan KA Reddy, William Huang arxiv

Voice activity detection (VAD) plays a vital role in enabling applications such as speech recognition. We analyze the impact of window size on the accuracy of three VAD algorithms: Silero, WebRTC, and Root Mean Square (RMS) across a set of diverse real-world digital audio streams. We additionally explore the use of hysteresis on top of each VAD output. Our results offer practical references for optimizing VAD systems. Silero significantly outperforms WebRTC and RMS, and hysteresis provides a benefit for WebRTC.

📄 PDF Abstract BibTeX arXiv:2601.17270

Code (0)

등록된 구현이 없습니다.

Tasks

Speech RecognitionActivity Detection

Similar Papers 제목 키워드 기반

On the Parameter Estimation of Sinusoidal Models for Speech and Audio Signals

2024-01-02 · George P. Kafentzis

In this paper, we examine the parameter estimation performance of three well-known sinusoidal models for speech and audio. The first one is the standard Sinusoidal Model (SM), which is based on the Fast Fourier Transform…

parameter estimationResynthesis

Vocal Prognostic Digital Biomarkers in Monitoring Chronic Heart Failure: A Longitudinal Observational Study

2026-03-31 · Fan Wu, Matthias P. Nägele, Daryush D. Mehta, Elgar Fleisch 외 arxiv

Objective: This study aimed to evaluate which voice features can predict health deterioration in patients with chronic HF. Background: Heart failure (HF) is a chronic condition with progressive deterioration and acute de…

EEG-Derived Voice Signature for Attended Speaker Detection

2023-08-28 · Hongxu Zhu, Siqi Cai, Yidi Jiang, Qiquan Zhang 외

\textit{Objective:} Conventional EEG-based auditory attention detection (AAD) is achieved by comparing the time-varying speech stimuli and the elicited EEG signals. However, in order to obtain reliable correlation values…

EEG

ASR data augmentation in low-resource settings using cross-lingual multi-speaker TTS and cross-lingual voice conversion

2022-03-29 · Edresson Casanova, Christopher Shulby, Alexander Korolev, Arnaldo Candido Junior 외

We explore cross-lingual multi-speaker speech synthesis and cross-lingual voice conversion applied to data augmentation for automatic speech recognition (ASR) systems in low/medium-resource scenarios. Through extensive e…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Data Augmentationspeech-recognition+3

Accurate Detection of Wake Word Start and End Using a CNN

2020-08-09 · Christin Jose, Yuriy Mishchenko, Thibaud Senechal, Anish Shah 외

Small footprint embedded devices require keyword spotters (KWS) with small model size and detection latency for enabling voice assistants. Such a keyword is often referred to as \textit{wake word} as it is used to wake u…