paper-with-me

Papers

SeqAfford: Sequential 3D Affordance Reasoning via Multimodal Large Language Model

2024-12-02 · CVPR 2025 1 · Chunlin Yu, Hanqing Wang, Ye Shi, Haoyang Luo, Sibei Yang, Jingyi Yu, Jingya Wang

3D affordance segmentation aims to link human instructions to touchable regions of 3D objects for embodied manipulations. Existing efforts typically adhere to single-object, single-affordance paradigms, where each affordance type or explicit instruction strictly corresponds to a specific affordance region and are unable to handle long-horizon tasks. Such a paradigm cannot actively reason about complex user intentions that often imply sequential affordances. In this paper, we introduce the Sequential 3D Affordance Reasoning task, which extends the traditional paradigm by reasoning from cumbersome user intentions and then decomposing them into a series of segmentation maps. Toward this, we construct the first instruction-based affordance segmentation benchmark that includes reasoning over both single and sequential affordances, comprising 180K instruction-point cloud pairs. Based on the benchmark, we propose our model, SeqAfford, to unlock the 3D multi-modal large language model with additional affordance segmentation abilities, which ensures reasoning with world knowledge and fine-grained affordance grounding in a cohesive framework. We further introduce a multi-granular language-point integration module to endow 3D dense prediction. Extensive experimental evaluations show that our model excels over well-established methods and exhibits open-world generalization with sequential reasoning abilities.

📄 PDF Abstract BibTeX arXiv:2412.01550

Code (0)

등록된 구현이 없습니다.

Tasks

Language ModelingLanguage ModellingLarge Language ModelMultimodal Large Language ModelSegmentationWorld Knowledge

Similar Papers 제목 키워드 기반

SeqAffordSplat: Scene-level Sequential Affordance Reasoning on 3D Gaussian Splatting

2025-07-31 · Di Li, Jie Feng, Jiahao Chen, Weisheng Dong 외 arxiv

3D affordance reasoning, the task of associating human instructions with the functional regions of 3D objects, is a critical capability for embodied agents. Current methods based on 3D Gaussian Splatting (3DGS) are funda…

Affordance-R1: Reinforcement Learning for Generalizable Affordance Reasoning in Multimodal Large Language Model

2025-08-08 · Hanqing Wang, Shaoyang Wang, Yiming Zhong, Zemin Yang 외 arxiv

Affordance grounding focuses on predicting the specific regions of objects that are associated with the actions to be performed by robots. It plays a vital role in the fields of human-robot interaction, human-object inte…

Zero-shot GeneralizationReinforcement Learning

A3R: Agentic Affordance Reasoning via Cross-Dimensional Evidence in 3D Gaussian Scenes

2026-04-02 · Di Li, Jie Feng, Guanbin Li, Ronghua Shang 외 arxiv

Affordance reasoning in 3D Gaussian scenes aims to identify the region that supports the action specified by a given text instruction in complex environments. Existing methods typically cast this problem as one-shot pred…

Decision Making

VideoAfford: Grounding 3D Affordance from Human-Object-Interaction Videos via Multimodal Large Language Model

2026-02-10 · Hanqing Wang, Mingyu Liu, Xiaoyu Chen, Chengwei MA 외 arxiv

3D affordance grounding aims to highlight the actionable regions on 3D objects, which is crucial for robotic manipulation. Previous research primarily focused on learning affordance knowledge from static cues such as lan…

Action UnderstandingPoint Clouds

Egocentric Instruction-oriented Affordance Prediction via Large Multimodal Model

2025-08-25 · Bokai Ji, Jie Gu, Xiaokang Ma, Chu Tang 외 arxiv

Affordance is crucial for intelligent robots in the context of object manipulation. In this paper, we argue that affordance should be task-/instruction-dependent, which is overlooked by many previous works. That is, diff…