paper-with-me

홈 › Papers

Multimedia Generative Script Learning for Task Planning

2022-08-25 · Qingyun Wang, Manling Li, Hou Pong Chan, Lifu Huang, Julia Hockenmaier, Girish Chowdhary, Heng Ji

Goal-oriented generative script learning aims to generate subsequent steps to reach a particular goal, which is an essential task to assist robots or humans in performing stereotypical activities. An important aspect of this process is the ability to capture historical states visually, which provides detailed information that is not covered by text and will guide subsequent steps. Therefore, we propose a new task, Multimedia Generative Script Learning, to generate subsequent steps by tracking historical states in both text and vision modalities, as well as presenting the first benchmark containing 5,652 tasks and 79,089 multimedia steps. This task is challenging in three aspects: the multimedia challenge of capturing the visual states in images, the induction challenge of performing unseen tasks, and the diversity challenge of covering different information in individual steps. We propose to encode visual state changes through a selective multimedia encoder to address the multimedia challenge, transfer knowledge from previously observed tasks using a retrieval-augmented decoder to overcome the induction challenge, and further present distinct information at each step by optimizing a diversity-oriented contrastive learning objective. We define metrics to evaluate both generation and inductive quality. Experiment results demonstrate that our approach significantly outperforms strong baselines.

📄 PDF Abstract BibTeX arXiv:2208.12306

Code (1)

EagleW/Multimedia-Generative-Script-Learning-for-Task-Planning 공식 구현 pytorch

Tasks

Contrastive LearningDecoderDescriptiveDiversityMultimedia Generative Script Learningmultimodal generationRetrievalTask Planning

Methods 이 논문이 사용한 방법론

Contrastive Learning 설명 없음

Similar Papers 제목 키워드 기반

Non-Sequential Graph Script Induction via Multimedia Grounding

2023-05-27 · Yu Zhou, Sha Li, Manling Li, Xudong Lin 외

Online resources such as WikiHow compile a wide range of scripts for performing everyday tasks, which can assist models in learning to reason about procedures. However, the scripts are always presented in a linear manner…

STIMONT: A core ontology for multimedia stimuli description

2014-01-10 · Marko Horvat, Nikola Bogunović, Krešimir Ćosić

Affective multimedia documents such as images, sounds or videos elicit emotional responses in exposed human subjects. These stimuli are stored in affective multimedia databases and successfully used for a wide variety of…

Retrieval

Learning and Inferring Movement with Deep Generative Model

2018-05-18 · Mingxuan Jing, Xiaojian Ma, Fuchun Sun, Huaping Liu

Learning and inference movement is a very challenging problem due to its high dimensionality and dependency to varied environments or tasks. In this paper, we propose an effective probabilistic method for learning and in…

modelMotion Planning

A Novel Approach to Multimedia Ontology Engineering for Automated Reasoning over Audiovisual LOD Datasets

2016-08-26 · Leslie F. Sikos

Multimedia reasoning, which is suitable for, among others, multimedia content analysis and high-level video scene interpretation, relies on the formal and comprehensive conceptualization of the represented knowledge doma…

Multimedia Verification Through Multi-Agent Deep Research Multimodal Large Language Models

2025-07-06 · Huy Hoan Le, Van Sy Thinh Nguyen, Thi Le Chi Dang, Vo Thanh Khang Nguyen 외 arxiv

This paper presents our submission to the ACMMM25 - Grand Challenge on Multimedia Verification. We developed a multi-agent verification system that combines Multimodal Large Language Models (MLLMs) with specialized verif…

Information Extraction