paper-with-me

Papers

MTPChat: A Multimodal Time-Aware Persona Dataset for Conversational Agents

2025-02-09 · Wanqi Yang, Yanda Li, Meng Fang, Ling Chen

Understanding temporal dynamics is critical for conversational agents, enabling effective content analysis and informed decision-making. However, time-aware datasets, particularly for persona-grounded conversations, are still limited, which narrows their scope and diminishes their complexity. To address this gap, we introduce MTPChat, a multimodal, time-aware persona dialogue dataset that integrates linguistic, visual, and temporal elements within dialogue and persona memory. Leveraging MTPChat, we propose two time-sensitive tasks: Temporal Next Response Prediction (TNRP) and Temporal Grounding Memory Prediction (TGMP), both designed to assess a model's ability to understand implicit temporal cues and dynamic interactions. Additionally, we present an innovative framework featuring an adaptive temporal module to effectively integrate multimodal streams and capture temporal dependencies. Experimental results validate the challenges posed by MTPChat and demonstrate the effectiveness of our framework in multimodal time-sensitive scenarios.

📄 PDF Abstract BibTeX arXiv:2502.05887

Code (0)

등록된 구현이 없습니다.

Tasks

Decision Making

Similar Papers 제목 키워드 기반

Personality-aware Human-centric Multimodal Reasoning: A New Task, Dataset and Baselines

2023-04-05 · Yaochen Zhu, Xiangqing Shen, Rui Xia

Personality traits, emotions, and beliefs shape individuals' behavioral choices and decision-making processes. However, for one thing, the affective computing community normally focused on predicting personality traits b…

Decision MakingMultimodal Reasoning

X-OmniClaw Technical Report: A Unified Mobile Agent for Multimodal Understanding and Interaction

2026-05-07 · Xiaoming Ren, Ru Zhen, Chao Li, Yang Song 외 arxiv

Inspired by the development of OpenClaw, there is a growing demand for mobile-based personal agents capable of handling complex and intuitive interactions. In this technical report, we introduce X-OmniClaw, a unified mob…

Multimodal Review Generation with Privacy and Fairness Awareness

2020-12-01 · COLING 2020 8 · Xuan-Son Vu, Thanh-Son Nguyen, Duc-Trong Le, Lili Jiang

Users express their opinions towards entities (e.g., restaurants) via online reviews which can be in diverse forms such as text, ratings, and images. Modeling reviews are advantageous for user behavior understanding whic…

FairnessReview GenerationSentiment AnalysisWord Embeddings

PHORECAST: Enabling AI Understanding of Public Health Outreach Across Populations

2025-10-02 · Rifaa Qadri, Anh Nhat Nhu, Swati Ramnath, Laura Yu Zheng 외 arxiv

Understanding how diverse individuals and communities respond to persuasive messaging holds significant potential for advancing personalized and socially aware machine learning. While Large Vision and Language Models (VL…

Region-Level Context-Aware Multimodal Understanding

2025-08-17 · Hongliang Wei, Xianqi Zhang, Xingtao Wang, Xiaopeng Fan 외 arxiv

Despite significant progress, existing research on Multimodal Large Language Models (MLLMs) mainly focuses on general visual understanding, overlooking the ability to integrate textual context associated with objects for…