paper-with-me

Papers

Microphone Array Geometry Independent Multi-Talker Distant ASR: NTT System for the DASR Task of the CHiME-8 Challenge

2025-02-14 · Naoyuki Kamo, Naohiro Tawara, Atsushi Ando, Takatomo Kano, Hiroshi Sato, Rintaro Ikeshita, Takafumi Moriya, Shota Horiguchi, Kohei Matsuura, Atsunori Ogawa, Alexis Plaquet, Takanori Ashihara, Tsubasa Ochiai, Masato Mimura, Marc Delcroix, Tomohiro Nakatani, Taichi Asami, Shoko Araki

In this paper, we introduce a multi-talker distant automatic speech recognition (DASR) system we designed for the DASR task 1 of the CHiME-8 challenge. Our system performs speaker counting, diarization, and ASR. It handles various recording conditions, from diner parties to professional meetings and from two to eight speakers. We perform diarization first, followed by speech enhancement, and then ASR as the challenge baseline. However, we introduced several key refinements. First, we derived a powerful speaker diarization relying on end-to-end speaker diarization with vector clustering (EEND-VC), multi-channel speaker counting using enhanced embeddings from EEND-VC, and target-speaker voice activity detection (TS-VAD). For speech enhancement, we introduced a novel microphone selection rule to better select the most relevant microphones among the distributed microphones and investigated improvements to beamforming. Finally, for ASR, we developed several models exploiting Whisper and WavLM speech foundation models. We present the results we submitted to the challenge and updated results we obtained afterward. Our strongest system achieves a 63% relative macro tcpWER improvement over the baseline and outperforms the challenge best results on the NOTSOFAR-1 meeting evaluation data among geometry-independent systems.

📄 PDF Abstract BibTeX arXiv:2502.09859

Code (0)

등록된 구현이 없습니다.

Tasks

Action DetectionActivity DetectionAutomatic Speech Recognitionspeaker-diarizationSpeaker DiarizationSpeech Enhancementspeech-recognitionSpeech Recognition

Similar Papers 제목 키워드 기반

VarArray Meets t-SOT: Advancing the State of the Art of Streaming Distant Conversational Speech Recognition

2022-09-12 · Naoyuki Kanda, Jian Wu, Xiaofei Wang, Zhuo Chen 외

This paper presents a novel streaming automatic speech recognition (ASR) framework for multi-talker overlapping speech captured by a distant microphone array with an arbitrary geometry. Our framework, named t-SOT-VA, cap…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition+1

Continuous Speech Separation with Ad Hoc Microphone Arrays

2021-03-03 · Dongmei Wang, Takuya Yoshioka, Zhuo Chen, Xiaofei Wang 외

Speech separation has been shown effective for multi-talker speech recognition. Under the ad hoc microphone array setup where the array consists of spatially distributed asynchronous microphones, additional challenges mu…

speech-recognitionSpeech RecognitionSpeech Separation

Multi-microphone Complex Spectral Mapping for Utterance-wise and Continuous Speech Separation

2020-10-04 · Zhong-Qiu Wang, Peidong Wang, DeLiang Wang

We propose multi-microphone complex spectral mapping, a simple way of applying deep learning for time-varying non-linear beamforming, for speaker separation in reverberant conditions. We aim at both speaker separation an…

Speaker SeparationSpeech Separation

Audio Inputs for Active Speaker Detection and Localization via Microphone Array

2023-07-27 · Davide Berghi, Philip J. B. Jackson

This study considers the problem of detecting and locating an active talker's horizontal position from multichannel audio captured by a microphone array. We refer to this as active speaker detection and localization (ASD…

Active Speaker Detection

UniX-Encoder: A Universal $X$-Channel Speech Encoder for Ad-Hoc Microphone Array Speech Processing

2023-10-25 · Zili Huang, Yiwen Shao, Shi-Xiong Zhang, Dong Yu

The speech field is evolving to solve more challenging scenarios, such as multi-channel recordings with multiple simultaneous talkers. Given the many types of microphone setups out there, we present the UniX-Encoder. It'…

speaker-diarizationSpeaker DiarizationSpeaker Recognitionspeech-recognition+1