paper-with-me

홈 › Papers

BEAR: A Video Dataset For Fine-grained Behaviors Recognition Oriented with Action and Environment Factors

2025-03-26 · Chengyang Hu, Yuduo Chen, Lizhuang Ma

Behavior recognition is an important task in video representation learning. An essential aspect pertains to effective feature learning conducive to behavior recognition. Recently, researchers have started to study fine-grained behavior recognition, which provides similar behaviors and encourages the model to concern with more details of behaviors with effective features for distinction. However, previous fine-grained behaviors limited themselves to controlling partial information to be similar, leading to an unfair and not comprehensive evaluation of existing works. In this work, we develop a new video fine-grained behavior dataset, named BEAR, which provides fine-grained (i.e. similar) behaviors that uniquely focus on two primary factors defining behavior: Environment and Action. It includes two fine-grained behavior protocols including Fine-grained Behavior with Similar Environments and Fine-grained Behavior with Similar Actions as well as multiple sub-protocols as different scenarios. Furthermore, with this new dataset, we conduct multiple experiments with different behavior recognition models. Our research primarily explores the impact of input modality, a critical element in studying the environmental and action-based aspects of behavior recognition. Our experimental results yield intriguing insights that have substantial implications for further research endeavors.

📄 PDF Abstract BibTeX arXiv:2503.20209

Code (0)

등록된 구현이 없습니다.

Tasks

Representation Learning

Methods 이 논문이 사용한 방법론

Focus 설명 없음

Similar Papers 제목 키워드 기반

Fine-Grained Egocentric Hand-Object Segmentation: Dataset, Model, and Applications

2022-08-07 · Lingzhi Zhang, Shenghao Zhou, Simon Stent, Jianbo Shi

Egocentric videos offer fine-grained information for high-fidelity modeling of human behaviors. Hands and interacting objects are one crucial aspect of understanding a viewer's behaviors and intentions. We provide a labe…

Activity RecognitionData AugmentationObjectSegmentation+2

Fine-grained Video Attractiveness Prediction Using Multimodal Deep Learning on a Large Real-world Dataset

2018-04-04 · Xinpeng Chen, Jingyuan Chen, Lin Ma, Jian Yao 외

Nowadays, billions of videos are online ready to be viewed and shared. Among an enormous volume of videos, some popular ones are widely viewed by online users while the majority attract little attention. Furthermore, wit…

Multimodal Deep LearningPrediction

Weakly-Supervised Temporal Action Detection for Fine-Grained Videos with Hierarchical Atomic Actions

2022-07-24 · Zhi Li, Lu He, Huijuan Xu

Action understanding has evolved into the era of fine granularity, as most human behaviors in real life have only minor differences. To detect these fine-grained actions accurately in a label-efficient way, we tackle the…

Action DetectionAction UnderstandingFine-Grained Action DetectionWeakly Supervised Action Localization

WTS: A Pedestrian-Centric Traffic Video Dataset for Fine-grained Spatial-Temporal Understanding

2024-07-22 · Quan Kong, Yuki Kawana, Rajat Saini, Ashutosh Kumar 외

In this paper, we address the challenge of fine-grained video event understanding in traffic scenarios, vital for autonomous driving and safety. Traditional datasets focus on driver or vehicle behavior, often neglecting …

Autonomous Driving

Dissecting Embodied Abilities in Multimodal Language Models through Skill-level Evaluation and Diagnosis

2025-10-09 · Yu Qi, Haibo Zhao, Ziyu Guo, Siyuan Ma 외 arxiv

Understanding the capability bottlenecks of embodied multimodal large language models (MLLMs) is crucial for improving embodied agents. However, existing embodied benchmarks mainly focus on task-level evaluation and fail…

Spatial Reasoning