paper-with-me

Papers

Understanding Multimodal Procedural Knowledge by Sequencing Multimodal Instructional Manuals

2021-10-16 · ACL 2022 5 · Te-Lin Wu, Alex Spangher, Pegah Alipoormolabashi, Marjorie Freedman, Ralph Weischedel, Nanyun Peng

The ability to sequence unordered events is an essential skill to comprehend and reason about real world task procedures, which often requires thorough understanding of temporal common sense and multimodal information, as these procedures are often communicated through a combination of texts and images. Such capability is essential for applications such as sequential task planning and multi-source instruction summarization. While humans are capable of reasoning about and sequencing unordered multimodal procedural instructions, whether current machine learning models have such essential capability is still an open question. In this work, we benchmark models' capability of reasoning over and sequencing unordered multimodal instructions by curating datasets from popular online instructional manuals and collecting comprehensive human annotations. We find models not only perform significantly worse than humans but also seem incapable of efficiently utilizing the multimodal information. To improve machines' performance on multimodal event sequencing, we propose sequentiality-aware pretraining techniques that exploit the sequential alignment properties of both texts and images, resulting in > 5% significant improvements.

📄 PDF Abstract BibTeX arXiv:2110.08486

Code (0)

등록된 구현이 없습니다.

Tasks

Common Sense ReasoningOpen-Ended Question AnsweringTask Planning

Similar Papers 제목 키워드 기반

Understanding Multimodal Procedural Knowledge by Sequencing Multimodal Instructional Manuals

2021-11-16 · ACL ARR November 2021 11 · Anonymous

The ability to sequence unordered events is evidence of comprehension and reasoning about real world tasks/procedures, and is essential for applications such as task planning and multi-source instruction summarization. I…

Common Sense ReasoningOpen-Ended Question AnsweringTask Planning

TAMA: Tool-Augmented Multimodal Agent for Procedural Activity Understanding

2025-09-30 · Kimihiro Hasegawa, Wiradee Imrattanatrai, Masaki Asada, Ken Fukuda 외 arxiv

Procedural activity assistants potentially support humans in a variety of settings, from our daily lives, e.g., cooking or assembling flat-pack furniture, to professional situations, e.g., manufacturing or biological exp…

Multimodal Reasoning

ProMQA: Question Answering Dataset for Multimodal Procedural Activity Understanding

2024-10-29 · Kimihiro Hasegawa, Wiradee Imrattanatrai, Zhi-Qi Cheng, Masaki Asada 외

Multimodal systems have great potential to assist humans in procedural activities, where people follow instructions to achieve their goals. Despite diverse application scenarios, systems are typically evaluated on tradit…

Action RecognitionAction SegmentationQuestion AnsweringTemporal Action Segmentation

RecipeQA: A Challenge Dataset for Multimodal Comprehension of Cooking Recipes

2018-09-04 · EMNLP 2018 10 · Semih Yagcioglu, Aykut Erdem, Erkut Erdem, Nazli Ikizler-Cinbis

Understanding and reasoning about cooking recipes is a fruitful research direction towards enabling machines to interpret procedural text. In this work, we introduce RecipeQA, a dataset for multimodal comprehension of co…

Reading Comprehension

MMSkills: Towards Multimodal Skills for General Visual Agents

2026-05-13 · Kangning Zhang, Shuai Shao, Qingyao Li, Jianghao Lin 외 arxiv

Reusable skills have become a core substrate for improving agent capabilities, yet most existing skill packages encode reusable behavior primarily as textual prompts, executable code, or learned routines. For visual agen…

Visual GroundingDecision Making