paper-with-me

Papers

Optimizing Speech Multi-View Feature Fusion through Conditional Computation

2025-01-14 · Weiqiao Shan, Yuhao Zhang, Yuchen Han, Bei Li, Xiaofeng Zhao, Yuang Li, Min Zhang, Hao Yang, Tong Xiao, Jingbo Zhu

Recent advancements have highlighted the efficacy of self-supervised learning (SSL) features in various speech-related tasks, providing lightweight and versatile multi-view speech representations. However, our study reveals that while SSL features expedite model convergence, they conflict with traditional spectral features like FBanks in terms of update directions. In response, we propose a novel generalized feature fusion framework grounded in conditional computation, featuring a gradient-sensitive gating network and a multi-stage dropout strategy. This framework mitigates feature conflicts and bolsters model robustness to multi-view input features. By integrating SSL and spectral features, our approach accelerates convergence and maintains performance on par with spectral models across multiple speech translation tasks on the MUSTC dataset.

📄 PDF Abstract BibTeX arXiv:2501.08057

Code (1)

shanweiqiao/gsgn 공식 구현

Tasks

Self-Supervised Learning

Methods 이 논문이 사용한 방법론

Dropout Dropout is a regularization technique for neural networks that drops a unit (along with connections) at training time with a specified probability $p$ (a common value is…

Similar Papers 제목 키워드 기반

Combining Multiple Views for Visual Speech Recognition

2017-10-19 · Marina Zimmermann, Mostafa Mehdipour Ghazi, Hazim Kemal Ekenel, Jean-Philippe Thiran

Visual speech recognition is a challenging research problem with a particular practical application of aiding audio speech recognition in noisy scenarios. Multiple camera setups can be beneficial for the visual speech re…

Sentencespeech-recognitionSpeech RecognitionVisual Speech Recognition

Attention Is Not Always the Answer: Optimizing Voice Activity Detection with Simple Feature Fusion

2025-06-02 · Kumud Tripathi, Chowdam Venkata Kumar, Pankaj Wasnik

Voice Activity Detection (VAD) plays a key role in speech processing, often utilizing hand-crafted or neural features. This study examines the effectiveness of Mel-Frequency Cepstral Coefficients (MFCCs) and pre-trained …

Action DetectionActivity DetectionComputational Efficiency

MDDM: A Multi-view Discriminative Enhanced Diffusion-based Model for Speech Enhancement

2025-05-19 · Nan Xu, Zhaolong Huang, Xiaonan Zhi

With the development of deep learning, speech enhancement has been greatly optimized in terms of speech quality. Previous methods typically focus on the discriminative supervised learning or generative modeling, which te…

Speech Enhancement

LoRA-Tuned Large Language Models for Dementia Detection via Multi-View Speech-Derived Features

2026-06-26 · Jonghyeon Park, Olivier Jiyoun Jung, Myungwoo Oh arxiv

Early detection of dementia enables timely intervention, and reflecting cognitive impairment, spontaneous speech offers a non-invasive screening modality. Conventional approaches often focus on a single representational …

Speech Recognition

Speech Emotion Recognition with Global-Aware Fusion on Multi-scale Feature Representation

2022-04-12 · Wenjing Zhu, Xiang Li

Speech Emotion Recognition (SER) is a fundamental task to predict the emotion label from speech data. Recent works mostly focus on using convolutional neural networks~(CNNs) to learn local attention map on fixed-scale fe…

Emotion RecognitionSpeech Emotion Recognition