paper-with-me

Papers

Stark: Social Long-Term Multi-Modal Conversation with Persona Commonsense Knowledge

2024-07-04 · Young-Jun Lee, Dokyong Lee, Junyoung Youn, Kyeongjin Oh, Byungsoo Ko, Jonghwan Hyeon, Ho-Jin Choi

Humans share a wide variety of images related to their personal experiences within conversations via instant messaging tools. However, existing works focus on (1) image-sharing behavior in singular sessions, leading to limited long-term social interaction, and (2) a lack of personalized image-sharing behavior. In this work, we introduce Stark, a large-scale long-term multi-modal conversation dataset that covers a wide range of social personas in a multi-modality format, time intervals, and images. To construct Stark automatically, we propose a novel multi-modal contextualization framework, Mcu, that generates long-term multi-modal dialogue distilled from ChatGPT and our proposed Plan-and-Execute image aligner. Using our Stark, we train a multi-modal conversation model, Ultron 7B, which demonstrates impressive visual imagination ability. Furthermore, we demonstrate the effectiveness of our dataset in human evaluation. We make our source code and dataset publicly available.

📄 PDF Abstract BibTeX arXiv:2407.03958

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

Focus 설명 없음

Similar Papers 제목 키워드 기반

Evaluating and Modeling Social Intelligence: A Comparative Study of Human and AI Capabilities

2024-05-20 · Junqi Wang, Chunhui Zhang, Jiapeng Li, Yuxi Ma 외

Facing the current debate on whether Large Language Models (LLMs) attain near-human intelligence levels (Mitchell & Krakauer, 2023; Bubeck et al., 2023; Kosinski, 2023; Shiffrin & Mitchell, 2023; Ullman, 2023), the curre…

Bayesian InferenceZero-Shot Learning

HourVideo: 1-Hour Video-Language Understanding

2024-11-07 · Keshigeyan Chandrasegaran, Agrim Gupta, Lea M. Hadzic, Taran Kota 외

We present HourVideo, a benchmark dataset for hour-long video-language understanding. Our dataset consists of a novel task suite comprising summarization, perception (recall, tracking), visual reasoning (spatial, tempora…

BenchmarkingcounterfactualMultiple-choiceRetrieval+1

MIMo: A Multi-Modal Infant Model for Studying Cognitive Development

2023-12-07 · Dominik Mattern, Pierre Schumacher, Francisco M. López, Marcel C. Raabe 외

Human intelligence and human consciousness emerge gradually during the process of cognitive development. Understanding this development is an essential aspect of understanding the human mind and may facilitate the constr…

Overabundant Information and Learning Traps

2018-05-21 · Annie Liang, Xiaosheng Mu

We develop a model of social learning from overabundant information: Short-lived agents sequentially choose from a large set of (flexibly correlated) information sources for prediction of an unknown state. Signal realiza…

OLViT: Multi-Modal State Tracking via Attention-Based Embeddings for Video-Grounded Dialog

2024-02-20 · Adnen Abdessaied, Manuel von Hochmeister, Andreas Bulling

We present the Object Language Video Transformer (OLViT) - a novel model for video dialog operating over a multi-modal attention-based dialog state tracker. Existing video dialog models struggle with questions requiring …

ObjectObject TrackingResponse GenerationTemporal Localization