paper-with-me

홈 › Papers

LSTM-based Whisper Detection

2018-09-20 · Zeynab Raeesy, Kellen Gillespie, Zhenpei Yang, Chengyuan Ma, Thomas Drugman, Jiacheng Gu, Roland Maas, Ariya Rastrow, Björn Hoffmeister

This article presents a whisper speech detector in the far-field domain. The proposed system consists of a long-short term memory (LSTM) neural network trained on log-filterbank energy (LFBE) acoustic features. This model is trained and evaluated on recordings of human interactions with voice-controlled, far-field devices in whisper and normal phonation modes. We compare multiple inference approaches for utterance-level classification by examining trajectories of the LSTM posteriors. In addition, we engineer a set of features based on the signal characteristics inherent to whisper speech, and evaluate their effectiveness in further separating whisper from normal speech. A benchmarking of these features using multilayer perceptrons (MLP) and LSTMs suggests that the proposed features, in combination with LFBE features, can help us further improve our classifiers. We prove that, with enough data, the LSTM model is indeed as capable of learning whisper characteristics from LFBE features alone compared to a simpler MLP model that uses both LFBE and features engineered for separating whisper and normal speech. In addition, we prove that the LSTM classifiers accuracy can be further improved with the incorporation of the proposed engineered features.

📄 PDF Abstract BibTeX arXiv:1809.07832

Code (0)

등록된 구현이 없습니다.

Tasks

Benchmarking

Methods 이 논문이 사용한 방법론

Sigmoid Activation 설명 없음
Tanh Activation 설명 없음
LSTM An LSTM is a type of recurrent neural network that addresses the vanishing gradient problem in vanilla…

Similar Papers 제목 키워드 기반

Identifying False Content and Hate Speech in Sinhala YouTube Videos by Analyzing the Audio

2024-01-30 · W. A. K. M. Wickramaarachchi, Sameeri Sathsara Subasinghe, K. K. Rashani Tharushika Wijerathna, A. Sahashra Udani Athukorala 외

YouTube faces a global crisis with the dissemination of false information and hate speech. To counter these issues, YouTube has implemented strict rules against uploading content that includes false information or promot…

Hate Speech DetectionMisinformationtext-classificationText Classification+1

Simul-Whisper: Attention-Guided Streaming Whisper with Truncation Detection

2024-06-14 · Haoyu Wang, Guoqiang Hu, Guodong Lin, Wei-Qiang Zhang 외

As a robust and large-scale multilingual speech recognition model, Whisper has demonstrated impressive results in many low-resource and out-of-distribution scenarios. However, its encoder-decoder structure hinders its ap…

Decoderspeech-recognitionSpeech Recognition

Deepfake Detection of Singing Voices With Whisper Encodings

2025-01-31 · Falguni Sharma, Priyanka Gupta

The deepfake generation of singing vocals is a concerning issue for artists in the music industry. In this work, we propose a singing voice deepfake detection (SVDD) system, which uses noise-variant encodings of open-AI'…

DeepFake DetectionFace Swapping

Deep Learning-Driven Multimodal Detection and Movement Analysis of Objects in Culinary

2025-08-21 · Tahoshin Alam Ishat, Mohammad Abdul Qayum arxiv

This is a research exploring existing models and fine tuning them to combine a YOLOv8 segmentation model, a LSTM model trained on hand point motion sequence and a ASR (whisper-base) to extract enough data for a LLM (Tiny…

Improved DeepFake Detection Using Whisper Features

2023-06-02 · Piotr Kawa, Marcin Plata, Michał Czuba, Piotr Szymański 외

With a recent influx of voice generation methods, the threat introduced by audio DeepFake (DF) is ever-increasing. Several different detection methods have been presented as a countermeasure. Many methods are based on so…

Automatic Speech RecognitionDeepFake DetectionFace Swappingspeech-recognition+1