paper-with-me

홈 › Papers

JRDB: A Dataset and Benchmark of Egocentric Robot Visual Perception of Humans in Built Environments

2019-10-25 · Roberto Martín-Martín, Mihir Patel, Hamid Rezatofighi, Abhijeet Shenoi, JunYoung Gwak, Eric Frankel, Amir Sadeghian, Silvio Savarese

We present JRDB, a novel egocentric dataset collected from our social mobile manipulator JackRabbot. The dataset includes 64 minutes of annotated multimodal sensor data including stereo cylindrical 360$^\circ$ RGB video at 15 fps, 3D point clouds from two Velodyne 16 Lidars, line 3D point clouds from two Sick Lidars, audio signal, RGB-D video at 30 fps, 360$^\circ$ spherical image from a fisheye camera and encoder values from the robot's wheels. Our dataset incorporates data from traditionally underrepresented scenes such as indoor environments and pedestrian areas, all from the ego-perspective of the robot, both stationary and navigating. The dataset has been annotated with over 2.3 million bounding boxes spread over 5 individual cameras and 1.8 million associated 3D cuboids around all people in the scenes totaling over 3500 time consistent trajectories. Together with our dataset and the annotations, we launch a benchmark and metrics for 2D and 3D person detection and tracking. With this dataset, which we plan on extending with further types of annotation in the future, we hope to provide a new source of data and a test-bench for research in the areas of egocentric robot vision, autonomous navigation, and all perceptual tasks around social robotics in human environments.

📄 PDF Abstract BibTeX arXiv:1910.11792

Code (1)

StanfordVL/JRMOT_ROS 공식 구현 pytorch

Tasks

Autonomous NavigationHuman Detection

Similar Papers 제목 키워드 기반

JRDB-Reasoning: A Difficulty-Graded Benchmark for Visual Reasoning in Robotics

2025-08-14 · Simindokht Jahangard, Mehrzad Mohammadi, Yi Shen, Zhixi Cai 외 arxiv

Recent advances in Vision-Language Models (VLMs) and large language models (LLMs) have greatly enhanced visual reasoning, a key capability for embodied AI agents like robots. However, existing visual reasoning benchmarks…

Visual Reasoning

JRDB-PanoTrack: An Open-world Panoptic Segmentation and Tracking Robotic Dataset in Crowded Human Environments

2024-04-02 · CVPR 2024 1 · Duy-Tho Le, Chenhui Gou, Stavya Datta, Hengcan Shi 외

Autonomous robot systems have attracted increasing research attention in recent years, where environment understanding is a crucial step for robot navigation, human-robot interaction, and decision. Real-world robot syste…

Decision MakingPanoptic SegmentationRobot Navigation

JRDB-Pose: A Large-scale Dataset for Multi-Person Pose Estimation and Tracking

2022-10-20 · CVPR 2023 1 · Edward Vendrow, Duy Tho Le, Jianfei Cai, Hamid Rezatofighi

Autonomous robotic systems operating in human environments must understand their surroundings to make accurate and safe decisions. In crowded human scenes with close-up human-robot interaction and robot navigation, a dee…

DiversityMulti-Person Pose EstimationMulti-Person Pose Estimation and TrackingPose Estimation+2

JRDB-Social: A Multifaceted Robotic Dataset for Understanding of Context and Dynamics of Human Interactions Within Social Groups

2024-04-06 · CVPR 2024 1 · Simindokht Jahangard, Zhixi Cai, Shiki Wen, Hamid Rezatofighi

Understanding human social behaviour is crucial in computer vision and robotics. Micro-level observations like individual actions fall short, necessitating a comprehensive approach that considers individual behaviour, in…

JRDB-Pose3D: A Multi-person 3D Human Pose and Shape Estimation Dataset for Robotics

2026-02-03 · Sandika Biswas, Kian Izadpanah, Hamid Rezatofighi arxiv

Real-world scenes are inherently crowded. Hence, estimating 3D poses of all nearby humans, tracking their movements over time, and understanding their activities within social and environmental contexts are essential for…

3D human pose and shape estimation3D Human Pose EstimationAutonomous DrivingRobot Navigation