paper-with-me

홈 › Papers

M-BEST-RQ: A Multi-Channel Speech Foundation Model for Smart Glasses

2024-09-17 · Yufeng Yang, Desh Raj, Ju Lin, Niko Moritz, Junteng Jia, Gil Keren, Egor Lakomkin, Yiteng Huang, Jacob Donley, Jay Mahadeokar, Ozlem Kalinli

The growing popularity of multi-channel wearable devices, such as smart glasses, has led to a surge of applications such as targeted speech recognition and enhanced hearing. However, current approaches to solve these tasks use independently trained models, which may not benefit from large amounts of unlabeled data. In this paper, we propose M-BEST-RQ, the first multi-channel speech foundation model for smart glasses, which is designed to leverage large-scale self-supervised learning (SSL) in an array-geometry agnostic approach. While prior work on multi-channel speech SSL only evaluated on simulated settings, we curate a suite of real downstream tasks to evaluate our model, namely (i) conversational automatic speech recognition (ASR), (ii) spherical active source localization, and (iii) glasses wearer voice activity detection, which are sourced from the MMCSG and EasyCom datasets. We show that a general-purpose M-BEST-RQ encoder is able to match or surpass supervised models across all tasks. For the conversational ASR task in particular, using only 8 hours of labeled speech, our model outperforms a supervised ASR baseline that is trained on 2000 hours of labeled data, which demonstrates the effectiveness of our approach.

📄 PDF Abstract BibTeX arXiv:2409.11494

Code (0)

등록된 구현이 없습니다.

Tasks

Action DetectionActivity DetectionAutomatic Speech RecognitionAutomatic Speech Recognition (ASR)Self-Supervised Learningspeech-recognitionSpeech Recognition

Similar Papers 제목 키워드 기반

Equipping LLM with Directional Multi-Talker Speech Understanding Capabilities

2026-02-06 · Ju Lin, Jing Pan, Ruizhi Li, Ming Sun 외 arxiv

Recent studies have demonstrated that prompting large language models (LLM) with audio encodings enables effective speech understanding capabilities. However, most speech LLMs are trained on single-channel, single-talker…

Speech Recognition

AGADIR: Towards Array-Geometry Agnostic Directional Speech Recognition

2024-01-18 · Ju Lin, Niko Moritz, Yiteng Huang, Ruiming Xie 외

Wearable devices like smart glasses are approaching the compute capability to seamlessly generate real-time closed captions for live conversations. We build on our recently introduced directional Automatic Speech Recogni…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

Identification of TV Channel Watching from Smart Meter Data Using Energy Disaggregation

2020-07-01 · Pascal A. Schirmer, Iosif Mporas, Akbar Sheikh-Akbari

Smart meters are used to measure the energy consumption of households. Specifically, within the energy consumption task smart meter have been used for load forecasting, reduction of consumer bills as well as reduction of…

Load Forecasting

LibriheavyMix: A 20,000-Hour Dataset for Single-Channel Reverberant Multi-Talker Speech Separation, ASR and Speaker Diarization

2024-09-01 · Zengrui Jin, Yifan Yang, Mohan Shi, Wei Kang 외

The evolving speech processing landscape is increasingly focused on complex scenarios like meetings or cocktail parties with multiple simultaneous speakers and far-field conditions. Existing methodologies for addressing …

speaker-diarizationSpeaker DiarizationSpeech Separation

MIMO-SPEECH: End-to-End Multi-Channel Multi-Speaker Speech Recognition

2019-10-15 · Xuankai Chang, Wangyou Zhang, Yanmin Qian, Jonathan Le Roux 외

Recently, the end-to-end approach has proven its efficacy in monaural multi-speaker speech recognition. However, high word error rates (WERs) still prevent these systems from being used in practical applications. On the …

speech-recognitionSpeech RecognitionSpeech Separation