paper-with-me

홈 › Papers

JRDB-Act: A Large-scale Dataset for Spatio-temporal Action, Social Group and Activity Detection

2021-06-16 · CVPR 2022 1 · Mahsa Ehsanpour, Fatemeh Saleh, Silvio Savarese, Ian Reid, Hamid Rezatofighi

The availability of large-scale video action understanding datasets has facilitated advances in the interpretation of visual scenes containing people. However, learning to recognise human actions and their social interactions in an unconstrained real-world environment comprising numerous people, with potentially highly unbalanced and long-tailed distributed action labels from a stream of sensory data captured from a mobile robot platform remains a significant challenge, not least owing to the lack of a reflective large-scale dataset. In this paper, we introduce JRDB-Act, as an extension of the existing JRDB, which is captured by a social mobile manipulator and reflects a real distribution of human daily-life actions in a university campus environment. JRDB-Act has been densely annotated with atomic actions, comprises over 2.8M action labels, constituting a large-scale spatio-temporal action detection dataset. Each human bounding box is labeled with one pose-based action label and multiple~(optional) interaction-based action labels. Moreover JRDB-Act provides social group annotation, conducive to the task of grouping individuals based on their interactions in the scene to infer their social activities~(common activities in each social group). Each annotated label in JRDB-Act is tagged with the annotators' confidence level which contributes to the development of reliable evaluation strategies. In order to demonstrate how one can effectively utilise such annotations, we develop an end-to-end trainable pipeline to learn and infer these tasks, i.e. individual action and social group detection. The data and the evaluation code is publicly available at https://jrdb.erc.monash.edu/.

📄 PDF Abstract BibTeX arXiv:2106.08827

Code (0)

등록된 구현이 없습니다.

Tasks

Action DetectionAction UnderstandingActivity Detection

Similar Papers 제목 키워드 기반

SoGAR: Self-supervised Spatiotemporal Attention-based Social Group Activity Recognition

2023-04-27 · Naga VS Raviteja Chappa, Pha Nguyen, Alexander H Nelson, Han-Seok Seo 외

This paper introduces a novel approach to Social Group Activity Recognition (SoGAR) using Self-supervised Transformers network that can effectively utilize unlabeled video data. To extract spatio-temporal information, we…

Activity RecognitionGroup Activity Recognition

JRDB-Pose: A Large-scale Dataset for Multi-Person Pose Estimation and Tracking

2022-10-20 · CVPR 2023 1 · Edward Vendrow, Duy Tho Le, Jianfei Cai, Hamid Rezatofighi

Autonomous robotic systems operating in human environments must understand their surroundings to make accurate and safe decisions. In crowded human scenes with close-up human-robot interaction and robot navigation, a dee…

DiversityMulti-Person Pose EstimationMulti-Person Pose Estimation and TrackingPose Estimation+2

PatchTraj: Unified Time-Frequency Representation Learning via Dynamic Patches for Trajectory Prediction

2025-07-25 · Yanghong Liu, Xingping Dong, Ming Li, Weixing Zhang 외 arxiv

Pedestrian trajectory prediction is crucial for autonomous driving and robotics. While existing point-based and grid-based methods expose two main limitations: insufficiently modeling human motion dynamics, as they fail …

Representation LearningTrajectory PredictionAutonomous Driving

Spatio-Temporal Proximity-Aware Dual-Path Model for Panoramic Activity Recognition

2024-03-21 · Sumin Lee, Yooseung Wang, Sangmin Woo, Changick Kim

Panoramic Activity Recognition (PAR) seeks to identify diverse human activities across different scales, from individual actions to social group and global activities in crowded panoramic scenes. PAR presents two major c…

Activity Recognition

MPT-PAR:Mix-Parameters Transformer for Panoramic Activity Recognition

2024-08-01 · Wenqing Gan, Yan Sun, Feiran Liu, Xiangfeng Luo

The objective of the panoramic activity recognition task is to identify behaviors at various granularities within crowded and complex environments, encompassing individual actions, social group activities, and global act…

Activity RecognitionRepresentation Learning