paper-with-me

Papers

PAVLM: Advancing Point Cloud based Affordance Understanding Via Vision-Language Model

2024-10-15 · Shang-Ching Liu, Van Nhiem Tran, Wenkai Chen, Wei-Lun Cheng, Yen-Lin Huang, I-Bin Liao, Yung-Hui Li, Jianwei Zhang

Affordance understanding, the task of identifying actionable regions on 3D objects, plays a vital role in allowing robotic systems to engage with and operate within the physical world. Although Visual Language Models (VLMs) have excelled in high-level reasoning and long-horizon planning for robotic manipulation, they still fall short in grasping the nuanced physical properties required for effective human-robot interaction. In this paper, we introduce PAVLM (Point cloud Affordance Vision-Language Model), an innovative framework that utilizes the extensive multimodal knowledge embedded in pre-trained language models to enhance 3D affordance understanding of point cloud. PAVLM integrates a geometric-guided propagation module with hidden embeddings from large language models (LLMs) to enrich visual semantics. On the language side, we prompt Llama-3.1 models to generate refined context-aware text, augmenting the instructional input with deeper semantic cues. Experimental results on the 3D-AffordanceNet benchmark demonstrate that PAVLM outperforms baseline methods for both full and partial point clouds, particularly excelling in its generalization to novel open-world affordance tasks of 3D objects. For more information, visit our project site: pavlm-source.github.io.

📄 PDF Abstract BibTeX arXiv:2410.11564

Code (0)

등록된 구현이 없습니다.

Tasks

Language ModelingLanguage Modelling

Similar Papers 제목 키워드 기반

3D AffordanceNet: A Benchmark for Visual Object Affordance Understanding

2021-03-30 · CVPR 2021 1 · Shengheng Deng, Xun Xu, Chaozheng Wu, Ke Chen 외

The ability to understand the ways to interact with objects from visual cues, a.k.a. visual affordance, is essential to vision-guided robotic research. This involves categorizing, segmenting and reasoning of visual affor…

Affordance DetectionBenchmarkingObject

Open-Vocabulary Affordance Detection in 3D Point Clouds

2023-03-04 · Toan Nguyen, Minh Nhat Vu, An Vuong, Dzung Nguyen 외

Affordance detection is a challenging problem with a wide variety of robotic applications. Traditional affordance detection methods are limited to a predefined set of affordance labels, hence potentially restricting the …

Affordance Detection

Task-Aware 3D Affordance Segmentation via 2D Guidance and Geometric Refinement

2025-11-12 · Lian He, Meng Liu, Qilang Ye, Yu Zhou 외 arxiv

Understanding 3D scene-level affordances from natural language instructions is essential for enabling embodied agents to interact meaningfully in complex environments. However, this task remains challenging due to the ne…

Affordance DetectionPoint Clouds

TB-HSU: Hierarchical 3D Scene Understanding with Contextual Affordances

2024-12-07 · Wenting Xu, Viorela Ila, Luping Zhou, Craig T. Jin

The concept of function and affordance is a critical aspect of 3D scene understanding and supports task-oriented objectives. In this work, we develop a model that learns to structure and vary functional affordance across…

Multi-Task LearningObjectScene Understanding

Interpretable Affordance Detection on 3D Point Clouds with Probabilistic Prototypes

2025-04-25 · Maximilian Xiling Li, Korbinian Rudolf, Nils Blank, Rudolf Lioutikov

Robotic agents need to understand how to interact with objects in their environment, both autonomously and during human-robot interactions. Affordance detection on 3D point clouds, which identifies object regions that al…

Affordance DetectionDecision Making