paper-with-me

Papers

RIO: A Benchmark for Reasoning Intention-Oriented Objects in Open Environments

2023-09-26 · NeurIPS 2023 11

Intention-oriented object detection aims to detect desired objects based on specific intentions or requirements. For instance, when we desire to "lie down and rest", we instinctively seek out a suitable option such as a "bed" or a "sofa" that can fulfill our needs. Previous work in this area is limited either by the number of intention descriptions or by the affordance vocabulary available for intention objects. These limitations make it challenging to handle intentions in open environments effectively. To facilitate this research, we construct a comprehensive dataset called Reasoning Intention-Oriented Objects (RIO). In particular, RIO is specifically designed to incorporate diverse real-world scenarios and a wide range of object categories. It offers the following key features: 1) intention descriptions in RIO are represented as natural sentences rather than a mere word or verb phrase, making them more practical and meaningful; 2) the intention descriptions are contextually relevant to the scene, enabling a broader range of potential functionalities associated with the objects; 3) the dataset comprises a total of 40,214 images and 130,585 intention-object pairs. With the proposed RIO, we evaluate the ability of some existing models to reason intention-oriented objects in open environments.Submission Number: 57

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Intention Reasoning Network for Multi-Domain End-to-end Task-Oriented Dialogue

2021-11-01 · EMNLP 2021 11 · Zhiyuan Ma, Jianjun Li, Zezheng Zhang, GuoHui Li 외

Recent years has witnessed the remarkable success in end-to-end task-oriented dialog system, especially when incorporating external knowledge information. However, the quality of most existing models’ generated response …

Task-oriented grasping for dexterous robots using postural synergies and reinforcement learning

2026-02-24 · Dimitrios Dimou, José Santos-Victor, Plinio Moreno arxiv

In this paper, we address the problem of task-oriented grasping for humanoid robots, emphasizing the need to align with human social norms and task-specific objectives. Existing methods, employ a variety of open-loop and…

Reinforcement Learning

DROGON: A Trajectory Prediction Model based on Intention-Conditioned Behavior Reasoning

2019-07-31 · Chiho Choi, Srikanth Malla, Abhishek Patil, Joon Hee Choi

We propose a Deep RObust Goal-Oriented trajectory prediction Network (DROGON) for accurate vehicle trajectory prediction by considering behavioral intentions of vehicles in traffic scenes. Our main insight is that the be…

Pedestrian Trajectory PredictionPredictionTrajectory Prediction

Visual Intention Grounding for Egocentric Assistants

2025-04-18 · Pengzhan Sun, Junbin Xiao, Tze Ho Elden Tse, Yicong Li 외

Visual grounding associates textual descriptions with objects in an image. Conventional methods target third-person image inputs and named object queries. In applications such as AI assistants, the perspective shifts -- …

ObjectVisual Grounding

SeqAfford: Sequential 3D Affordance Reasoning via Multimodal Large Language Model

2024-12-02 · CVPR 2025 1 · Chunlin Yu, Hanqing Wang, Ye Shi, Haoyang Luo 외

3D affordance segmentation aims to link human instructions to touchable regions of 3D objects for embodied manipulations. Existing efforts typically adhere to single-object, single-affordance paradigms, where each afford…

Language ModelingLanguage ModellingLarge Language ModelMultimodal Large Language Model+2