paper-with-me

Papers

Understanding Multimodal Procedural Knowledge by Sequencing Multimodal Instructional Manuals

2021-11-16 · ACL ARR November 2021 11 · Anonymous

The ability to sequence unordered events is evidence of comprehension and reasoning about real world tasks/procedures, and is essential for applications such as task planning and multi-source instruction summarization. It often requires thorough understanding of temporal common sense and multimodal information, since these procedures are often conveyed by a combination of texts and images. While humans are capable of reasoning about and sequencing unordered procedural instructions, the extent to which the current machine learning methods possess such a capability is still an open question. In this work, we benchmark models' capability of reasoning over and sequencing unordered multimodal instructions by curating datasets from online instructional manuals and collecting comprehensive human annotations. We find current state-of-the-art models not only perform significantly worse than humans but also seem incapable of efficiently utilizing multimodal information. To improve machines' performance on multimodal event sequencing, we propose sequence-aware pretraining techniques exploiting the sequential alignment properties of both texts and images, resulting in > 5% improvements on perfect match ratio.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Common Sense ReasoningOpen-Ended Question AnsweringTask Planning

Similar Papers 제목 키워드 기반

Understanding Multimodal Procedural Knowledge by Sequencing Multimodal Instructional Manuals

2021-10-16 · ACL 2022 5 · Te-Lin Wu, Alex Spangher, Pegah Alipoormolabashi, Marjorie Freedman 외

The ability to sequence unordered events is an essential skill to comprehend and reason about real world task procedures, which often requires thorough understanding of temporal common sense and multimodal information, a…

Common Sense ReasoningOpen-Ended Question AnsweringTask Planning

TAMA: Tool-Augmented Multimodal Agent for Procedural Activity Understanding

2025-09-30 · Kimihiro Hasegawa, Wiradee Imrattanatrai, Masaki Asada, Ken Fukuda 외 arxiv

Procedural activity assistants potentially support humans in a variety of settings, from our daily lives, e.g., cooking or assembling flat-pack furniture, to professional situations, e.g., manufacturing or biological exp…

Multimodal Reasoning

ProMQA: Question Answering Dataset for Multimodal Procedural Activity Understanding

2024-10-29 · Kimihiro Hasegawa, Wiradee Imrattanatrai, Zhi-Qi Cheng, Masaki Asada 외

Multimodal systems have great potential to assist humans in procedural activities, where people follow instructions to achieve their goals. Despite diverse application scenarios, systems are typically evaluated on tradit…

Action RecognitionAction SegmentationQuestion AnsweringTemporal Action Segmentation

RecipeQA: A Challenge Dataset for Multimodal Comprehension of Cooking Recipes

2018-09-04 · EMNLP 2018 10 · Semih Yagcioglu, Aykut Erdem, Erkut Erdem, Nazli Ikizler-Cinbis

Understanding and reasoning about cooking recipes is a fruitful research direction towards enabling machines to interpret procedural text. In this work, we introduce RecipeQA, a dataset for multimodal comprehension of co…

Reading Comprehension

MMSkills: Towards Multimodal Skills for General Visual Agents

2026-05-13 · Kangning Zhang, Shuai Shao, Qingyao Li, Jianghao Lin 외 arxiv

Reusable skills have become a core substrate for improving agent capabilities, yet most existing skill packages encode reusable behavior primarily as textual prompts, executable code, or learned routines. For visual agen…

Visual GroundingDecision Making