paper-with-me

홈 › Papers

Hybrid Attention based Multimodal Network for Spoken Language Classification

2018-08-01 · COLING 2018 8 · Yue Gu, Kangning Yang, Shiyu Fu, Shuhong Chen, Xinyu Li, Ivan Marsic

We examine the utility of linguistic content and vocal characteristics for multimodal deep learning in human spoken language understanding. We present a deep multimodal network with both feature attention and modality attention to classify utterance-level speech data. The proposed hybrid attention architecture helps the system focus on learning informative representations for both modality-specific feature extraction and model fusion. The experimental results show that our system achieves state-of-the-art or competitive results on three published multimodal datasets. We also demonstrated the effectiveness and generalization of our system on a medical speech dataset from an actual trauma scenario. Furthermore, we provided a detailed comparison and analysis of traditional approaches and deep learning methods on both feature extraction and fusion.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

ClassificationDeep LearningEmotion RecognitionGeneral ClassificationMultimodal Deep LearningSentiment AnalysisSpoken Language Understanding

Similar Papers 제목 키워드 기반

Deep Multimodal Learning for Emotion Recognition in Spoken Language

2018-02-22 · Yue Gu, Shuhong Chen, Ivan Marsic

In this paper, we present a novel deep multimodal framework to predict human emotions based on sentence-level spoken language. Our architecture has two distinctive characteristics. First, it extracts the high-level featu…

Emotion RecognitionSentence

ESPnet-ST-v2: Multipurpose Spoken Language Translation Toolkit

2023-04-10 · Brian Yan, Jiatong Shi, Yun Tang, Hirofumi Inaguma 외

ESPnet-ST-v2 is a revamp of the open-source ESPnet-ST toolkit necessitated by the broadening interests of the spoken language translation community. ESPnet-ST-v2 supports 1) offline speech-to-text translation (ST), 2) si…

BenchmarkingSimultaneous Speech-to-Text TranslationSpeech-to-Speech TranslationSpeech-to-Text+2

UAM: A Unified Attention-Mamba Backbone of Multimodal Framework for Tumor Cell Classification

2025-11-21 · Taixi Chen, Jingyun Chen, Nancy Guo arxiv

Inspired by the recent success of the Mamba architecture in vision and language domains, we introduce a Unified Attention-Mamba (UAM) backbone. Unlike previous hybrid approaches that integrate Attention and Mamba modules…

Tumor SegmentationImage Segmentation

Multimodal Speech Recognition for Language-Guided Embodied Agents

2023-02-27 · Allen Chang, Xiaoyuan Zhu, Aarav Monga, Seoho Ahn 외

Benchmarks for language-guided embodied agents typically assume text-based instructions, but deployed agents will encounter spoken instructions. While Automatic Speech Recognition (ASR) models can bridge the input gap, e…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

A novel multimodal dynamic fusion network for disfluency detection in spoken utterances

2022-11-27 · Sreyan Ghosh, Utkarsh Tyagi, Sonal Kumar, Manan Suri 외

Disfluency, though originating from human spoken utterances, is primarily studied as a uni-modal text-based Natural Language Processing (NLP) task. Based on early-fusion and self-attention-based multimodal interaction be…

multimodal interaction