paper-with-me

홈 › Papers

3M-TRANSFORMER: A Multi-Stage Multi-Stream Multimodal Transformer for Embodied Turn-Taking Prediction

2023-10-23 · Mehdi Fatan, Emanuele Mincato, Dimitra Pintzou, Mariella Dimiccoli

Predicting turn-taking in multiparty conversations has many practical applications in human-computer/robot interaction. However, the complexity of human communication makes it a challenging task. Recent advances have shown that synchronous multi-perspective egocentric data can significantly improve turn-taking prediction compared to asynchronous, single-perspective transcriptions. Building on this research, we propose a new multimodal transformer-based architecture for predicting turn-taking in embodied, synchronized multi-perspective data. Our experimental results on the recently introduced EgoCom dataset show a substantial performance improvement of up to 14.01% on average compared to existing baselines and alternative transformer-based approaches. The source code, and the pre-trained models of our 3M-Transformer will be available upon acceptance.

📄 PDF Abstract BibTeX arXiv:2310.14859

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

CoSMo: A Multimodal Transformer for Page Stream Segmentation in Comic Books

2025-07-14 · Marc Serra Ortega, Emanuele Vivoli, Artemis Llabrés, Dimosthenis Karatzas arxiv

This paper introduces CoSMo, a novel multimodal Transformer for Page Stream Segmentation (PSS) in comic books, a critical task for automated content understanding, as it is a necessary first stage for many downstream tas…

UniVL: A Unified Video and Language Pre-Training Model for Multimodal Understanding and Generation

2020-02-15 · Huaishao Luo, Lei Ji, Botian Shi, Haoyang Huang 외

With the recent success of the pre-training technique for NLP and image-linguistic tasks, some video-linguistic pre-training works are gradually developed to improve video-text related downstream tasks. However, most of …

Action SegmentationDecoderLanguage ModelingLanguage Modelling+2

Accountable Textual-Visual Chat Learns to Reject Human Instructions in Image Re-creation

2023-03-10 · Zhiwei Zhang, Yuliang Liu

The recent success of ChatGPT and GPT-4 has drawn widespread attention to multimodal dialogue systems. However, there is a lack of datasets in the academic community that can effectively evaluate the multimodal generatio…

Image Generationmultimodal generationText to Image GenerationText-to-Image Generation+1

StreaMulT: Streaming Multimodal Transformer for Heterogeneous and Arbitrary Long Sequential Data

2021-10-15 · Victor Pellegrain, Myriam Tami, Michel Batteux, Céline Hudelot

The increasing complexity of Industry 4.0 systems brings new challenges regarding predictive maintenance tasks such as fault detection and diagnosis. A corresponding and realistic setting includes multi-source data strea…

Fault DetectionMultimodal Sentiment AnalysisSentiment Analysis

CorMulT: A Semi-supervised Modality Correlation-aware Multimodal Transformer for Sentiment Analysis

2024-07-09 · Yangmin Li, Ruiqi Zhu, Wengen Li

Multimodal sentiment analysis is an active research area that combines multiple data modalities, e.g., text, image and audio, to analyze human emotions and benefits a variety of applications. Existing multimodal sentimen…

Contrastive LearningMultimodal Sentiment AnalysisPredictionSentiment Analysis