paper-with-me

홈 › Papers

ActPlan-1K: Benchmarking the Procedural Planning Ability of Visual Language Models in Household Activities

2024-10-04 · Ying Su, Zhan Ling, Haochen Shi, Jiayang Cheng, Yauwai Yim, Yangqiu Song

Large language models~(LLMs) have been adopted to process textual task description and accomplish procedural planning in embodied AI tasks because of their powerful reasoning ability. However, there is still lack of study on how vision language models~(VLMs) behave when multi-modal task inputs are considered. Counterfactual planning that evaluates the model's reasoning ability over alternative task situations are also under exploited. In order to evaluate the planning ability of both multi-modal and counterfactual aspects, we propose ActPlan-1K. ActPlan-1K is a multi-modal planning benchmark constructed based on ChatGPT and household activity simulator iGibson2. The benchmark consists of 153 activities and 1,187 instances. Each instance describing one activity has a natural language task description and multiple environment images from the simulator. The gold plan of each instance is action sequences over the objects in provided scenes. Both the correctness and commonsense satisfaction are evaluated on typical VLMs. It turns out that current VLMs are still struggling at generating human-level procedural plans for both normal activities and counterfactual activities. We further provide automatic evaluation metrics by finetuning over BLEURT model to facilitate future research on our benchmark.

📄 PDF Abstract BibTeX arXiv:2410.03907

Code (0)

등록된 구현이 없습니다.

Tasks

BenchmarkingcounterfactualCounterfactual Planning

Similar Papers 제목 키워드 기반

LLaPa: A Vision-Language Model Framework for Counterfactual-Aware Procedural Planning

2025-07-11 · Shibo Sun, Xue Li, Donglin Di, Mingjie Wei 외 arxiv

While large language models (LLMs) have advanced procedural planning for embodied AI systems through strong reasoning abilities, the integration of multimodal inputs and counterfactual reasoning remains underexplored. To…

GaLa: Hypergraph-Guided Visual Language Models for Procedural Planning

2026-04-19 · Kun Wang, Yiming Li, Mingcheng Qu, Aqiang Zhang 외 arxiv

Implicit spatial relations and deep semantic structures encoded in object attributes are crucial for procedural planning in embodied AI systems. However, existing approaches often over rely on the reasoning capabilities …

Contrastive Learning

InFactPlanner: Planning Sustainable Geo-Distributed LLM Data Centers

2026-08-13 · Nicoletta Tsiopani, Moysis Symeonides, George Pallis, Marios D. Dikaiakos arxiv

The rapid growth of LLM inference is shifting sustainability concerns from one-time training to continuous serving, where infrastructure decisions shape energy use, carbon emissions, water consumption, and service qualit…

Procedural Generalization by Planning with Self-Supervised World Models

2021-11-02 · ICLR 2022 4 · Ankesh Anand, Jacob Walker, Yazhe Li, Eszter Vértes 외

One of the key promises of model-based reinforcement learning is the ability to generalize using an internal model of the world to make predictions in novel environments and tasks. However, the generalization ability of …

BenchmarkingMeta-LearningModel-based Reinforcement LearningRepresentation Learning

VG-TVP: Multimodal Procedural Planning via Visually Grounded Text-Video Prompting

2024-12-16 · Muhammet Furkan Ilaslan, Ali Koksal, Kevin Qinhong Lin, Burak Satar 외

Large Language Model (LLM)-based agents have shown promise in procedural tasks, but the potential of multimodal instructions augmented by texts and videos to assist users remains under-explored. To address this gap, we p…

InformativenessLarge Language ModelText GenerationText-to-Video Generation+2