paper-with-me

홈 › Papers

From Multi-modal Property Dataset to Robot-centric Conceptual Knowledge About Household Objects

2019-06-26 · Madhura Thosar, Christian A. Mueller, Georg Jaeger, Johannes Schleiss, Narender Pulugu, Ravi Mallikarjun Chennaboina, Sai Vivek Jeevangekar, Andreas Birk, Max Pfingsthorn, Sebastian Zug

Tool-use applications in robotics require conceptual knowledge about objects for informed decision making and object interactions. State-of-the-art methods employ hand-crafted symbolic knowledge which is defined from a human perspective and grounded into sensory data afterwards. However, due to different sensing and acting capabilities of robots, their conceptual understanding of objects must be generated from a robot's perspective entirely, which asks for robot-centric conceptual knowledge about objects. With this goal in mind, this article motivates that such knowledge should be based on physical and functional properties of objects. Consequently, a selection of ten properties is defined and corresponding extraction methods are proposed. This multi-modal property extraction forms the basis on which our second contribution, a robot-centric knowledge generation is build on. It employs unsupervised clustering methods to transform numerical property data into symbols, and Bivariate Joint Frequency Distributions and Sample Proportion to generate conceptual knowledge about objects using the robot-centric symbols. A preliminary implementation of the proposed framework is employed to acquire a dataset comprising physical and functional property data of 110 houshold objects. This Robot-Centric dataSet (RoCS) is used to evaluate the framework regarding the property extraction methods, the semantics of the considered properties within the dataset and its usefulness in real-world applications such as tool substitution.

📄 PDF Abstract BibTeX arXiv:1906.11114

Code (0)

등록된 구현이 없습니다.

Tasks

ClusteringDecision Making

Similar Papers 제목 키워드 기반

Seeing Together: Multi-Robot Cooperative Egocentric Spatial Reasoning with Multimodal Large Language Models

2026-05-18 · Kunyu Peng, Zhikun Zhou, Kailun Yang, Di Wen 외 arxiv

Multimodal Large Language Models (MLLMs) have made substantial progress in egocentric video understanding, but their ability to reason cooperatively from multiple embodied viewpoints remains largely unexplored. We study …

Spatial Reasoning

Egocentric RGB+Depth Action Recognition in Industry-Like Settings

2023-09-25 · Jyoti Kini, Sarah Fleischer, Ishan Dave, Mubarak Shah

Action recognition from an egocentric viewpoint is a crucial perception task in robotics and enables a wide range of human-robot interactions. While most computer vision approaches prioritize the RGB camera, the Depth mo…

Action Recognition

Aria-NeRF: Multimodal Egocentric View Synthesis

2023-11-11 · Jiankai Sun, Jianing Qiu, Chuanyang Zheng, John Tucker 외

We seek to accelerate research in developing rich, multimodal scene models trained from egocentric data, based on differentiable volumetric ray-tracing inspired by Neural Radiance Fields (NeRFs). The construction of a Ne…

NeRF

Egocentric Human Trajectory Forecasting with a Wearable Camera and Multi-Modal Fusion

2021-11-01 · Jianing Qiu, Lipeng Chen, Xiao Gu, Frank P. -W. Lo 외

In this paper, we address the problem of forecasting the trajectory of an egocentric camera wearer (ego-person) in crowded spaces. The trajectory forecasting ability learned from the data of different camera wearers walk…

DecoderTrajectory Forecasting

J-CRe3: A Japanese Conversation Dataset for Real-world Reference Resolution

2024-03-28 · Nobuhiro Ueda, Hideko Habe, Yoko Matsui, Akishige Yuguchi 외

Understanding expressions that refer to the physical world is crucial for such human-assisting systems in the real world, as robots that must perform actions that are expected by users. In real-world reference resolution…