paper-with-me

홈 › Papers

Reproducible Multimodal Affordance Prediction

2026-08-18 · Tommaso Apicella, Alessio Xompero, Andrea Cavallaro arxiv

Affordance prediction is the identification of potential actions an agent can perform on a target object from multimodal inputs. Affordance prediction methods are difficult to evaluate and compare due to heterogeneous problem formulations, inconsistent dataset annotations, incomplete reporting of experimental protocols, and limited information about deployment conditions. These limitations challenge fair benchmarking and performance comparison. To promote transparency, we propose the Affordance Sheet, a documentation detailing task formulation with its input modalities, model architectures and training information, datasets, and experimental protocols. Affordance Sheets enable reproducible benchmarking and reliable evaluation of affordance models for real-world scenarios, including generalisation to novel conditions and human safety.

📄 PDF Abstract BibTeX arXiv:2608.18317

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Egocentric Instruction-oriented Affordance Prediction via Large Multimodal Model

2025-08-25 · Bokai Ji, Jie Gu, Xiaokang Ma, Chu Tang 외 arxiv

Affordance is crucial for intelligent robots in the context of object manipulation. In this paper, we argue that affordance should be task-/instruction-dependent, which is overlooked by many previous works. That is, diff…

Segmenting Object Affordances: Reproducibility and Sensitivity to Scale

2024-09-03 · Tommaso Apicella, Alessio Xompero, Paolo Gastaldo, Andrea Cavallaro

Visual affordance segmentation identifies image regions of an object an agent can interact with. Existing methods re-use and adapt learning-based architectures for semantic segmentation to the affordance segmentation tas…

ObjectSegmentationSemantic SegmentationSensitivity

RoboAfford++: A Generative AI-Enhanced Dataset for Multimodal Affordance Learning in Robotic Manipulation and Navigation

2025-11-16 · Xiaoshuai Hao, Yingbo Tang, Lingfeng Zhang, Yanbiao Ma 외 arxiv

Robotic manipulation and navigation are fundamental capabilities of embodied intelligence, enabling effective robot interactions with the physical world. Achieving these capabilities requires a cohesive understanding of …

Affordance RecognitionScene UnderstandingObject RecognitionQuestion Answering

SeqAfford: Sequential 3D Affordance Reasoning via Multimodal Large Language Model

2024-12-02 · CVPR 2025 1 · Chunlin Yu, Hanqing Wang, Ye Shi, Haoyang Luo 외

3D affordance segmentation aims to link human instructions to touchable regions of 3D objects for embodied manipulations. Existing efforts typically adhere to single-object, single-affordance paradigms, where each afford…

Language ModelingLanguage ModellingLarge Language ModelMultimodal Large Language Model+2

Text-driven object affordance for guiding grasp-type recognition in multimodal robot teaching

2021-02-27 · Naoki Wake, Daichi Saito, Kazuhiro Sasabuchi, Hideki Koike 외

This study investigates how text-driven object affordance, which provides prior knowledge about grasp types for each object, affects image-based grasp-type recognition in robot teaching. The researchers created labeled d…

Mixed RealityObjectVocal Bursts Type Prediction