paper-with-me

Papers

MMS-LLaMA: Efficient LLM-based Audio-Visual Speech Recognition with Minimal Multimodal Speech Tokens

2025-03-14 · Jeong Hun Yeo, Hyeongseop Rha, Se Jin Park, Yong Man Ro

Audio-Visual Speech Recognition (AVSR) achieves robust speech recognition in noisy environments by combining auditory and visual information. However, recent Large Language Model (LLM) based AVSR systems incur high computational costs due to the high temporal resolution of audio-visual speech processed by LLMs. In this work, we introduce an efficient multimodal speech LLM framework that minimizes token length while preserving essential linguistic content. Our approach employs an early AV-fusion module for streamlined feature integration, an audio-visual speech Q-Former that dynamically allocates tokens based on input duration, and a refined query allocation strategy with a speech rate predictor to adjust token allocation according to speaking speed of each audio sample. Extensive experiments on the LRS3 dataset show that our method achieves state-of-the-art performance with a WER of 0.72% while using only 3.5 tokens per second. Moreover, our approach not only reduces token usage by 86% compared to the previous multimodal speech LLM framework, but also improves computational efficiency by reducing FLOPs by 35.7%.

📄 PDF Abstract BibTeX arXiv:2503.11315

Code (1)

JeongHun0716/MMS-LLaMA 공식 구현 pytorch

Tasks

Audio-Visual Speech RecognitionComputational EfficiencyLanguage ModelingLanguage ModellingLarge Language ModelRobust Speech Recognitionspeech-recognitionSpeech RecognitionVisual Speech Recognition

Methods 이 논문이 사용한 방법론

SPEED The monocular depth estimation (MDE) is the task of estimating depth from a single frame. This information is an essential knowledge in many computer vision tasks such as scene…

Similar Papers 제목 키워드 기반

Large Language Models are Strong Audio-Visual Speech Recognition Learners

2024-09-18 · Umberto Cappellazzo, Minsu Kim, Honglie Chen, Pingchuan Ma 외

Multimodal large language models (MLLMs) have recently become a focal point of research due to their formidable multimodal understanding capabilities. For example, in the audio and speech domains, an LLM can be equipped …

Audio-Visual Speech RecognitionAutomatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognition+2

Adaptive Audio-Visual Speech Recognition via Matryoshka-Based Multimodal LLMs

2025-03-09 · Umberto Cappellazzo, Minsu Kim, Stavros Petridis

Audio-Visual Speech Recognition (AVSR) leverages both audio and visual modalities to enhance speech recognition robustness, particularly in noisy environments. Recent advancements in Large Language Models (LLMs) have dem…

Audio-Visual Speech RecognitionComputational EfficiencyRepresentation Learningspeech-recognition+2

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition

2025-02-11 · Sungnyun Kim, Kangwook Jang, Sangmin Bae, Sungwoo Cho 외

Audio-visual speech recognition (AVSR) has become critical for enhancing speech recognition in noisy environments by integrating both auditory and visual modalities. However, existing AVSR systems struggle to scale up wi…

Audio-Visual Speech RecognitionComputational EfficiencyMixture-of-ExpertsRobust Speech Recognition+3

Thinking in Directivity: Speech Large Language Model for Multi-Talker Directional Speech Recognition

2025-06-17 · Jiamin Xie, Ju Lin, Yiteng Huang, Tyler Vuong 외

Recent studies have demonstrated that prompting large language models (LLM) with audio encodings enables effective speech recognition capabilities. However, the ability of Speech LLMs to comprehend and process multi-chan…

Data AugmentationLanguage ModelingLanguage ModellingLarge Language Model+2

AV-CPL: Continuous Pseudo-Labeling for Audio-Visual Speech Recognition

2023-09-29 · Andrew Rouditchenko, Ronan Collobert, Tatiana Likhomanenko

Audio-visual speech contains synchronized audio and visual information that provides cross-modal supervision to learn representations for both automatic speech recognition (ASR) and visual speech recognition (VSR). We in…

Audio-Visual Speech RecognitionAutomatic Speech RecognitionAutomatic Speech Recognition (ASR)Pseudo Label+3