paper-with-me

홈 › Papers

A Study on Speech Assessment with Visual Cues

2025-06-11 · Shafique Ahmed, Ryandhimas E. Zezario, Nasir Saleem, Amir Hussain, Hsin-Min Wang, Yu Tsao

Non-intrusive assessment of speech quality and intelligibility is essential when clean reference signals are unavailable. In this work, we propose a multimodal framework that integrates audio features and visual cues to predict PESQ and STOI scores. It employs a dual-branch architecture, where spectral features are extracted using STFT, and visual embeddings are obtained via a visual encoder. These features are then fused and processed by a CNN-BLSTM with attention, followed by multi-task learning to simultaneously predict PESQ and STOI. Evaluations on the LRS3-TED dataset, augmented with noise from the DEMAND corpus, show that our model outperforms the audio-only baseline. Under seen noise conditions, it improves LCC by 9.61% (0.8397->0.9205) for PESQ and 11.47% (0.7403->0.8253) for STOI. These results highlight the effectiveness of incorporating visual cues in enhancing the accuracy of non-intrusive speech assessment.

📄 PDF Abstract BibTeX arXiv:2506.09549

Code (0)

등록된 구현이 없습니다.

Tasks

Multi-Task Learning

Methods 이 논문이 사용한 방법론

LCC Please enter a description about the method here

Similar Papers 제목 키워드 기반

Artificial Intelligence for Suicide Assessment using Audiovisual Cues: A Review

2022-01-22 · Sahraoui Dhelim, Liming Chen, Huansheng Ning, Chris Nugent

Death by suicide is the seventh leading death cause worldwide. The recent advancement in Artificial Intelligence (AI), specifically AI applications in image and voice processing, has created a promising opportunity to re…

The complementary roles of non-verbal cues for Robust Pronunciation Assessment

2023-09-14 · Yassine El Kheir, Shammur Absar Chowdhury, Ahmed Ali

Research on pronunciation assessment systems focuses on utilizing phonetic and phonological aspects of non-native (L2) speech, often neglecting the rich layer of information hidden within the non-verbal cues. In this stu…

On the Role of Visual Cues in Audiovisual Speech Enhancement

2020-04-25 · Zakaria Aldeneh, Anushree Prasanna Kumar, Barry-John Theobald, Erik Marchi 외

We present an introspection of an audiovisual speech enhancement model. In particular, we focus on interpreting how a neural audiovisual speech enhancement model uses visual cues to improve the quality of the target spee…

Self-Supervised LearningSpeech Enhancement

DAMOS: Learning Distortion-Aware Speech Quality Assessment through Explicit Distortion Localization

2026-08-21 · Naiyuan Li, Li Dong, Diqun Yan arxiv

Automatic speech quality assessment aims to predict Mean Opinion Scores (MOS) consistent with human subjective perception and is essential for evaluating speech generation, enhancement, and communication systems. For spe…

Muse: Multi-modal target speaker extraction with visual cues

2020-10-15 · Zexu Pan, Ruijie Tao, Chenglin Xu, Haizhou Li

Speaker extraction algorithm relies on the speech sample from the target speaker as the reference point to focus its attention. Such a reference speech is typically pre-recorded. On the other hand, the temporal synchroni…

Target Speaker Extraction