paper-with-me

Papers

Weakly-supervised Multi-task Learning for Multimodal Affect Recognition

2021-04-23 · Wenliang Dai, Samuel Cahyawijaya, Yejin Bang, Pascale Fung

Multimodal affect recognition constitutes an important aspect for enhancing interpersonal relationships in human-computer interaction. However, relevant data is hard to come by and notably costly to annotate, which poses a challenging barrier to build robust multimodal affect recognition systems. Models trained on these relatively small datasets tend to overfit and the improvement gained by using complex state-of-the-art models is marginal compared to simple baselines. Meanwhile, there are many different multimodal affect recognition datasets, though each may be small. In this paper, we propose to leverage these datasets using weakly-supervised multi-task learning to improve the generalization performance on each of them. Specifically, we explore three multimodal affect recognition tasks: 1) emotion recognition; 2) sentiment analysis; and 3) sarcasm recognition. Our experimental results show that multi-tasking can benefit all these tasks, achieving an improvement up to 2.9% accuracy and 3.3% F1-score. Furthermore, our method also helps to improve the stability of model performance. In addition, our analysis suggests that weak supervision can provide a comparable contribution to strong supervision if the tasks are highly correlated.

📄 PDF Abstract BibTeX arXiv:2104.11560

Code (0)

등록된 구현이 없습니다.

Tasks

Emotion RecognitionMulti-Task LearningSentiment Analysis

Similar Papers 제목 키워드 기반

Can audio-visual integration strengthen robustness under multimodal attacks?

2021-04-05 · CVPR 2021 1 · Yapeng Tian, Chenliang Xu

In this paper, we propose to make a systematic study on machines multisensory perception under attacks. We use the audio-visual event recognition task against multimodal adversarial attacks as a proxy to investigate the …

audio-visual learningVisual Localization

Weakly Supervised Multimodal Temporal Forgery Localization via Multitask Learning

2025-08-04 · Wenbo Xu, Wei Lu, Xiangyang Luo arxiv

The spread of Deepfake videos has caused a trust crisis and impaired social stability. Although numerous approaches have been proposed to address the challenges of Deepfake detection and localization, there is still a la…

Binary ClassificationDeepFake Detection

MAF: Multimodal Alignment Framework for Weakly-Supervised Phrase Grounding

2020-10-12 · EMNLP 2020 11 · Qinxin Wang, Hao Tan, Sheng Shen, Michael W. Mahoney 외

Phrase localization is a task that studies the mapping from textual phrases to regions of an image. Given difficulties in annotating phrase-to-object datasets at scale, we develop a Multimodal Alignment Framework (MAF) t…

Phrase Grounding

Learning weakly supervised multimodal phoneme embeddings

2017-04-23 · Rahma Chaabouni, Ewan Dunbar, Neil Zeghidour, Emmanuel Dupoux

Recent works have explored deep architectures for learning multimodal speech representation (e.g. audio and images, articulation and audio) in a supervised way. Here we investigate the role of combining different speech …

Multi-Task Learning

Weakly-supervised anomaly detection for multimodal data distributions

2024-06-13 · Xu Tan, Junqi Chen, Sylwan Rahardja, Jiawei Yang 외

Weakly-supervised anomaly detection can outperform existing unsupervised methods with the assistance of a very small number of labeled anomalies, which attracts increasing attention from researchers. However, existing we…

Anomaly DetectionSupervised Anomaly DetectionWeakly-supervised Anomaly Detection