paper-with-me

Papers

Tailor Versatile Multi-modal Learning for Multi-label Emotion Recognition

2022-01-15 · Yi Zhang, Mingyuan Chen, Jundong Shen, Chongjun Wang

Multi-modal Multi-label Emotion Recognition (MMER) aims to identify various human emotions from heterogeneous visual, audio and text modalities. Previous methods mainly focus on projecting multiple modalities into a common latent space and learning an identical representation for all labels, which neglects the diversity of each modality and fails to capture richer semantic information for each label from different perspectives. Besides, associated relationships of modalities and labels have not been fully exploited. In this paper, we propose versaTile multi-modAl learning for multI-labeL emOtion Recognition (TAILOR), aiming to refine multi-modal representations and enhance discriminative capacity of each label. Specifically, we design an adversarial multi-modal refinement module to sufficiently explore the commonality among different modalities and strengthen the diversity of each modality. To further exploit label-modal dependence, we devise a BERT-like cross-modal encoder to gradually fuse private and common modality representations in a granularity descent way, as well as a label-guided decoder to adaptively generate a tailored representation for each label with the guidance of label semantics. In addition, we conduct experiments on the benchmark MMER dataset CMU-MOSEI in both aligned and unaligned settings, which demonstrate the superiority of TAILOR over the state-of-the-arts. Code is available at https://github.com/kniter1/TAILOR.

📄 PDF Abstract BibTeX arXiv:2201.05834

Code (1)

kniter1/tailor 공식 구현 pytorch

Tasks

DecoderDiversityEmotion Recognition

Similar Papers 제목 키워드 기반

PowMix: A Versatile Regularizer for Multimodal Sentiment Analysis

2023-12-19 · Efthymios Georgiou, Yannis Avrithis, Alexandros Potamianos

Multimodal sentiment analysis (MSA) leverages heterogeneous data sources to interpret the complex nature of human sentiments. Despite significant progress in multimodal architecture design, the field lacks comprehensive …

Multimodal Sentiment AnalysisSentiment Analysis

The Multi-modality Cell Segmentation Challenge: Towards Universal Solutions

2023-08-10 · Jun Ma, Ronald Xie, Shamini Ayyadhury, Cheng Ge 외

Cell segmentation is a critical step for quantitative single-cell analysis in microscopy images. Existing cell segmentation methods are often tailored to specific modalities or require manual interventions to specify hyp…

Cell SegmentationSegmentation

SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding

2025-08-30 · Zhen Chen, Xingjian Luo, Kun Yuan, Jinlin Wu 외 arxiv

Surgical video understanding is crucial for facilitating Computer-Assisted Surgery (CAS) systems. Despite significant progress in existing studies, two major limitations persist, including inadequate visual content perce…

Video Reconstruction

The CAMOMILE Collaborative Annotation Platform for Multi-modal, Multi-lingual and Multi-media Documents

2016-05-01 · LREC 2016 5 · Johann Poignant, Mateusz Budnik, Herv{\'e} Bredin, Claude Barras 외

In this paper, we describe the organization and the implementation of the CAMOMILE collaborative annotation framework for multimodal, multimedia, multilingual (3M) data. Given the versatile nature of the analysis which c…

Active LearningManagement

WAVE: Learning Unified & Versatile Audio-Visual Embeddings with Multimodal LLM

2025-09-26 · Changli Tang, Qinfan Xiao, Ke Mei, Tianyi Wang 외 arxiv

While embeddings from multimodal large language models (LLMs) excel as general-purpose representations, their application to dynamic modalities like audio and video remains underexplored. We introduce WAVE (\textbf{u}nif…

Cross-Modal RetrievalQuestion Answering