paper-with-me

Papers

Leveraging TCN and Transformer for effective visual-audio fusion in continuous emotion recognition

2023-03-15 · Weiwei Zhou, Jiada Lu, Zhaolong Xiong, Weifeng Wang

Human emotion recognition plays an important role in human-computer interaction. In this paper, we present our approach to the Valence-Arousal (VA) Estimation Challenge, Expression (Expr) Classification Challenge, and Action Unit (AU) Detection Challenge of the 5th Workshop and Competition on Affective Behavior Analysis in-the-wild (ABAW). Specifically, we propose a novel multi-modal fusion model that leverages Temporal Convolutional Networks (TCN) and Transformer to enhance the performance of continuous emotion recognition. Our model aims to effectively integrate visual and audio information for improved accuracy in recognizing emotions. Our model outperforms the baseline and ranks 3 in the Expression Classification challenge.

📄 PDF Abstract BibTeX arXiv:2303.08356

Code (1)

upczww/abaw5 공식 구현 pytorch

Tasks

Emotion Recognition

Methods 이 논문이 사용한 방법론

Multi-Head Attention 설명 없음
Attention 설명 없음
Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
Adam 설명 없음
Position-Wise Feed-Forward Layer 설명 없음
Residual Connection 설명 없음
BPE Byte Pair Encoding, or BPE, is a subword segmentation algorithm that encodes rare and unknown words as sequences of subword units. The intuition is that various word…
Dense Connections Dense Connections, or Fully Connected Connections, are a type of layer in a deep neural network that use a linear operation where every input is connected to every output…

Similar Papers 제목 키워드 기반

Dynamic Multi-Target Fusion for Efficient Audio-Visual Navigation

2025-09-23 · Yinfeng Yu, Hailong Zhang, Meiling Zhu arxiv

Audiovisual embodied navigation enables robots to locate audio sources by dynamically integrating visual observations from onboard sensors with the auditory signals emitted by the target. The core challenge lies in effec…

Visual Navigation

Unveiling the Power of Audio-Visual Early Fusion Transformers with Dense Interactions through Masked Modeling

2023-12-02 · CVPR 2024 1 · Shentong Mo, Pedro Morgado

Humans possess a remarkable ability to integrate auditory and visual information, enabling a deeper understanding of the surrounding environment. This early fusion of audio and visual cues, demonstrated through cognitive…

Audio-Visual Speech Recognition based on Regulated Transformer and Spatio-Temporal Fusion Strategy for Driver Assistive Systems

2024-05-09 · Expert Systems with Applications 2024 5 · Dmitry Ryumin, Alexandr Axyonov, Elena Ryumina, Denis Ivanko 외

This article presents a research methodology for audio-visual speech recognition (AVSR) in driver assistive systems. These systems necessitate ongoing interaction with drivers while driving through voice control for safe…

Audio-Visual Speech RecognitionLipreadingLip Readingspeech-recognition+2

AVT: Audio-Video Transformer for Multimodal Action Recognition

2022-09-22 · Submitted to ICLR 2022 9 · Wentao Zhu, Jingru Yi, Kevin Hsu, Xiaohang Sun 외

Action recognition is an essential field for video understanding. To learn from heterogeneous data sources effectively, in this work, we propose a novel multimodal action recognition approach termed Audio-Video Transform…

Action RecognitionAudio ClassificationContrastive LearningMulti-modal Classification+1

Attentive Fusion Enhanced Audio-Visual Encoding for Transformer Based Robust Speech Recognition

2020-08-06 · Liangfa Wei, Jie Zhang, JunFeng Hou, Li-Rong Dai

Audio-visual information fusion enables a performance improvement in speech recognition performed in complex acoustic scenarios, e.g., noisy environments. It is required to explore an effective audio-visual fusion strate…

Robust Speech Recognitionspeech-recognitionSpeech Recognition