Pretraining on Interactions for Learning Grounded Affordance Representations
Lexical semantics and cognitive science point to affordances (i.e. the actions that objects support) as critical for understanding and representing nouns and verbs. However, study of these semantic features has not yet been integrated with the "foundation" models that currently dominate language representation research. We hypothesize that predictive modeling of object state over time will result in representations that encode object affordance information "for free". We train a neural network to predict objects' trajectories in a simulated interaction and show that our network's latent representations differentiate between both observed and unobserved affordances. We find that models trained using 3D simulations from our SPATIAL dataset outperform conventional 2D computer vision models trained on a similar task, and, on initial inspection, that differences between concepts correspond to expected features (e.g., roll entails rotation). Our results suggest a way in which modern deep learning approaches to grounded language learning can be integrated with traditional formal semantic notions of lexical representations.
Code (1)
Tasks
Grounded language learningSimilar Papers 제목 키워드 기반
Pretraining over Interactions for Learning Grounded Object Representations
Large language models have been criticized for their limited ability to reason about \textit{affordances} - the actions that can be performed on an object. It has been argued that to accomplish this, models need some for…
ObjectMulti-label affordance mapping from egocentric vision
Accurate affordance detection and segmentation with pixel precision is an important piece in many complex systems based on interactions, such as robots and assitive devices. We present a new approach to affordance percep…
Affordance DetectionSegmentationAffordance Transfer Learning for Human-Object Interaction Detection
Reasoning the human-object interactions (HOI) is essential for deeper scene understanding, while object affordances (or functionalities) are of great importance for human to discover unseen HOIs with novel objects. Inspi…
Affordance DetectionAffordance RecognitionHuman-Object Interaction Concept DiscoveryHuman-Object Interaction Detection+3Grounded Affordance from Exocentric View
Affordance grounding aims to locate objects' "action possibilities" regions, which is an essential step toward embodied intelligence. Due to the diversity of interactive affordance, the uniqueness of different individual…
DiversityHuman-Object Interaction DetectionObjectTransfer LearningVision-Guided Action: Enhancing 3D Human Motion Prediction with Gaze-informed Affordance in 3D Scenes
Recent advances in human motion prediction (HMP) have shifted focus from isolated motion data to integrating human-scene correlations. In particular, the latest methods leverage human gaze points, using their spatial…
Human motion predictionHuman-Object Interaction Detectionmotion prediction