paper-with-me

홈 › Papers

Beyond Voice Activity Detection: Hybrid Audio Segmentation for Direct Speech Translation

2021-04-23 · ICNLSP 2021 11 · Marco Gaido, Matteo Negri, Mauro Cettolo, Marco Turchi

The audio segmentation mismatch between training data and those seen at run-time is a major problem in direct speech translation. Indeed, while systems are usually trained on manually segmented corpora, in real use cases they are often presented with continuous audio requiring automatic (and sub-optimal) segmentation. After comparing existing techniques (VAD-based, fixed-length and hybrid segmentation methods), in this paper we propose enhanced hybrid solutions to produce better results without sacrificing latency. Through experiments on different domains and language pairs, we show that our methods outperform all the other techniques, reducing by at least 30% the gap between the traditional VAD-based approach and optimal manual segmentation.

📄 PDF Abstract BibTeX arXiv:2104.11710

Code (0)

등록된 구현이 없습니다.

Tasks

Action DetectionActivity DetectionSegmentationTranslation

Similar Papers 제목 키워드 기반

Kernel-based Sensor Fusion with Application to Audio-Visual Voice Activity Detection

2016-04-11 · David Dov, Ronen Talmon, Israel Cohen

In this paper, we address the problem of multiple view data fusion in the presence of noise and interferences. Recent studies have approached this problem using kernel methods, by relying particularly on a product of ker…

Action DetectionActivity DetectionSensor Fusion

Intel Labs at Ego4D Challenge 2022: A Better Baseline for Audio-Visual Diarization

2022-10-14 · Kyle Min

This report describes our approach for the Audio-Visual Diarization (AVD) task of the Ego4D Challenge 2022. Specifically, we present multiple technical improvements over the official baselines. First, we improve the dete…

Action DetectionActive Speaker DetectionActivity Detection

Speaker Embeddings as Individuality Proxy for Voice Stress Detection

2023-06-09 · Zihan Wu, Neil Scheidwasser-Clow, Karl El Hajal, Milos Cernak

Since the mental states of the speaker modulate speech, stress introduced by cognitive or physical loads could be detected in the voice. The existing voice stress detection benchmark has shown that the audio embeddings e…

Multi-Input Multi-Output Target-Speaker Voice Activity Detection For Unified, Flexible, and Robust Audio-Visual Speaker Diarization

2024-01-16 · Ming Cheng, Ming Li

Audio-visual learning has demonstrated promising results in many classical speech tasks (e.g., speech separation, automatic speech recognition, wake-word spotting). We believe that introducing visual modality will also b…

Action DetectionActivity Detectionaudio-visual learningAutomatic Speech Recognition+5

Multi-Stage Speaker Diarization for Noisy Classrooms

2025-05-16 · Ali Sartaz Khan, Tolulope Ogunremi, Ahmed Adel Attia, Dorottya Demszky

Speaker diarization, the process of identifying "who spoke when" in audio recordings, is essential for understanding classroom dynamics. However, classroom settings present distinct challenges, including poor recording q…

Action DetectionActivity DetectionAutomatic Speech RecognitionAutomatic Speech Recognition (ASR)+5