paper-with-me

Papers

TextMI: Textualize Multimodal Information for Integrating Non-verbal Cues in Pre-trained Language Models

2023-03-27 · Md Kamrul Hasan, Md Saiful Islam, Sangwu Lee, Wasifur Rahman, Iftekhar Naim, Mohammed Ibrahim Khan, Ehsan Hoque

Pre-trained large language models have recently achieved ground-breaking performance in a wide variety of language understanding tasks. However, the same model can not be applied to multimodal behavior understanding tasks (e.g., video sentiment/humor detection) unless non-verbal features (e.g., acoustic and visual) can be integrated with language. Jointly modeling multiple modalities significantly increases the model complexity, and makes the training process data-hungry. While an enormous amount of text data is available via the web, collecting large-scale multimodal behavioral video datasets is extremely expensive, both in terms of time and money. In this paper, we investigate whether large language models alone can successfully incorporate non-verbal information when they are presented in textual form. We present a way to convert the acoustic and visual information into corresponding textual descriptions and concatenate them with the spoken text. We feed this augmented input to a pre-trained BERT model and fine-tune it on three downstream multimodal tasks: sentiment, humor, and sarcasm detection. Our approach, TextMI, significantly reduces model complexity, adds interpretability to the model's decision, and can be applied for a diverse set of tasks while achieving superior (multimodal sarcasm detection) or near SOTA (multimodal sentiment analysis and multimodal humor detection) performance. We propose TextMI as a general, competitive baseline for multimodal behavioral analysis tasks, particularly in a low-resource setting.

📄 PDF Abstract BibTeX arXiv:2303.15430

Code (0)

등록된 구현이 없습니다.

Tasks

Humor DetectionMultimodal Sentiment AnalysisSarcasm DetectionSentiment Analysis

Methods 이 논문이 사용한 방법론

Refunds@Expedia|||How do I get a full refund from Expedia? “How do I get a full refund from Expedia? How do I get a full refund from Expedia? – Call ☎️ +1-(888) 829 (0881) or +1-805-330-4056 or +1-805-330-4056 for Quick Help &…
Multi-Head Attention 설명 없음
Attention 설명 없음
Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
Dense Connections Dense Connections, or Fully Connected Connections, are a type of layer in a deep neural network that use a linear operation where every input is connected to every output…
Layer Normalization Unlike batch normalization, Layer Normalization directly estimates the normalization statistics from the summed inputs…
Weight Decay 설명 없음
Adam 설명 없음

Similar Papers 제목 키워드 기반

WDMIR: Wavelet-Driven Multimodal Intent Recognition

2025-05-27 · Weiyin Gong, Kai Zhang, Yanghai Zhang, Qi Liu 외

Multimodal intent recognition (MIR) seeks to accurately interpret user intentions by integrating verbal and non-verbal information across video, audio and text modalities. While existing approaches prioritize text analys…

Intent RecognitionMultimodal Intent Recognition

MOTIF: Contextualized Images for Complex Words to Improve Human Reading

2022-06-01 · LREC 2022 6 · Xintong Wang, Florian Schneider, Özge Alacam, Prateek Chaudhury 외

MOTIF (MultimOdal ConTextualized Images For Language Learners) is a multimodal dataset that consists of 1125 comprehension texts retrieved from Wikipedia Simple Corpus. Allowing multimodal processing or enriching the con…

Reading Comprehension

ContextMix: A context-aware data augmentation method for industrial visual inspection systems

2024-01-18 · Hyungmin Kim, Donghun Kim, Pyunghwan Ahn, Sungho Suh 외

While deep neural networks have achieved remarkable performance, data augmentation has emerged as a crucial strategy to mitigate overfitting and enhance network performance. These techniques hold particular significance …

Data AugmentationObject Recognition

Improving Machine Reading Comprehension with Contextualized Commonsense Knowledge

2020-09-12 · ACL 2022 5 · Kai Sun, Dian Yu, Jianshu Chen, Dong Yu 외

In this paper, we aim to extract commonsense knowledge to improve machine reading comprehension. We propose to represent relations implicitly by situating structured knowledge in a context instead of relying on a pre-def…

Machine Reading ComprehensionReading Comprehension

Textualized and Feature-based Models for Compound Multimodal Emotion Recognition in the Wild

2024-07-17 · Nicolas Richet, Soufiane Belharbi, Haseeb Aslam, Meike Emilie Schadt 외

Systems for multimodal emotion recognition (ER) are commonly trained to extract features from different modalities (e.g., visual, audio, and textual) that are combined to predict individual basic emotions. However, compo…

Emotion RecognitionMultimodal Emotion Recognition