paper-with-me

홈 › Papers

Show and Guide: Instructional-Plan Grounded Vision and Language Model

2024-09-27 · Diogo Glória-Silva, David Semedo, João Magalhães

Guiding users through complex procedural plans is an inherently multimodal task in which having visually illustrated plan steps is crucial to deliver an effective plan guidance. However, existing works on plan-following language models (LMs) often are not capable of multimodal input and output. In this work, we present MM-PlanLLM, the first multimodal LLM designed to assist users in executing instructional tasks by leveraging both textual plans and visual information. Specifically, we bring cross-modality through two key tasks: Conversational Video Moment Retrieval, where the model retrieves relevant step-video segments based on user queries, and Visually-Informed Step Generation, where the model generates the next step in a plan, conditioned on an image of the user's current progress. MM-PlanLLM is trained using a novel multitask-multistage approach, designed to gradually expose the model to multimodal instructional-plans semantic layers, achieving strong performance on both multimodal and textual dialogue in a plan-grounded setting. Furthermore, we show that the model delivers cross-modal temporal and plan-structure representations aligned between textual plan steps and instructional video moments.

📄 PDF Abstract BibTeX arXiv:2409.19074

Code (1)

dmgcsilva/mmplanllm 공식 구현 pytorch

Tasks

Language ModelingLanguage ModellingMoment Retrieval

Similar Papers 제목 키워드 기반

VIGiA: Instructional Video Guidance via Dialogue Reasoning and Retrieval

2026-02-22 · Diogo Glória-Silva, David Semedo, João Maglhães arxiv

We introduce VIGiA, a novel multimodal dialogue model designed to understand and reason over complex, multi-step instructional video action plans. Unlike prior work which focuses mainly on text-only guidance, or treats v…

Event-Guided Procedure Planning from Instructional Videos with Text Supervision

2023-08-17 · ICCV 2023 1 · An-Lan Wang, Kun-Yu Lin, Jia-Run Du, Jingke Meng 외

In this work, we focus on the task of procedure planning from instructional videos with text supervision, where a model aims to predict an action sequence to transform the initial visual state into the goal visual state.…

TeachingCoach: A Fine-Tuned Scaffolding Chatbot for Instructional Guidance to Instructors

2026-03-18 · Isabel Molnar, Peiyu Li, Si Chen, Sugana Chawla 외 arxiv

Higher education instructors often lack timely and pedagogically grounded support, as scalable instructional guidance remains limited and existing tools rely on generic chatbot advice or non-scalable teaching center huma…

Dialogue Generation

Gondola: Grounded Vision Language Planning for Generalizable Robotic Manipulation

2025-06-12 · ShiZhe Chen, Ricardo Garcia, Paul Pacaud, Cordelia Schmid

Robotic manipulation faces a significant challenge in generalizing across unseen objects, environments and tasks specified by diverse language instructions. To improve generalization capabilities, recent research has inc…

Referring Expression

VEGGIE: Instructional Editing and Reasoning Video Concepts with Grounded Generation

2025-03-18 · Shoubin Yu, Difan Liu, Ziqiao Ma, Yicong Hong 외

Recent video diffusion models have enhanced video editing, but it remains challenging to handle instructional editing and diverse tasks (e.g., adding, removing, changing) within a unified framework. In this paper, we int…

Reasoning SegmentationVideo Editing