paper-with-me

홈 › Papers

DAP3D-Net: Where, What and How Actions Occur in Videos?

2016-02-10 · Li Liu, Yi Zhou, Ling Shao

Action parsing in videos with complex scenes is an interesting but challenging task in computer vision. In this paper, we propose a generic 3D convolutional neural network in a multi-task learning manner for effective Deep Action Parsing (DAP3D-Net) in videos. Particularly, in the training phase, action localization, classification and attributes learning can be jointly optimized on our appearancemotion data via DAP3D-Net. For an upcoming test video, we can describe each individual action in the video simultaneously as: Where the action occurs, What the action is and How the action is performed. To well demonstrate the effectiveness of the proposed DAP3D-Net, we also contribute a new Numerous-category Aligned Synthetic Action dataset, i.e., NASA, which consists of 200; 000 action clips of more than 300 categories and with 33 pre-defined action attributes in two hierarchical levels (i.e., low-level attributes of basic body part movements and high-level attributes related to action motion). We learn DAP3D-Net using the NASA dataset and then evaluate it on our collected Human Action Understanding (HAU) dataset. Experimental results show that our approach can accurately localize, categorize and describe multiple actions in realistic videos.

📄 PDF Abstract BibTeX arXiv:1602.03346

Code (0)

등록된 구현이 없습니다.

Tasks

Action LocalizationAction ParsingAction UnderstandingMulti-Task Learning

Similar Papers 제목 키워드 기반

Human Hands as Probes for Interactive Object Understanding

2021-12-16 · CVPR 2022 1 · Mohit Goyal, Sahil Modi, Rishabh Goyal, Saurabh Gupta

Interactive object understanding, or what we can do to objects and how is a long-standing goal of computer vision. In this paper, we tackle this problem through observation of human hands in in-the-wild egocentric videos…

Object

V-STaR: Benchmarking Video-LLMs on Video Spatio-Temporal Reasoning

2025-03-14 · Zixu Cheng, Jian Hu, Ziquan Liu, Chenyang Si 외

Human processes video reasoning in a sequential spatio-temporal reasoning logic, we first identify the relevant frames ("when") and then analyse the spatial relationships ("where") between key objects, and finally levera…

BenchmarkingRelational ReasoningVideo Understanding

When will you do what? - Anticipating Temporal Occurrences of Activities

2018-04-03 · CVPR 2018 6 · Yazan Abu Farha, Alexander Richard, Juergen Gall

Analyzing human actions in videos has gained increased attention recently. While most works focus on classifying and labeling observed video frames or anticipating the very recent future, making long-term predictions ove…

Weakly-supervised Visual Instrument-playing Action Detection in Videos

2018-05-05 · Jen-Yu Liu, Yi-Hsuan Yang, Shyh-Kang Jeng

Instrument playing is among the most common scenes in music-related videos, which represent nowadays one of the largest sources of online videos. In order to understand the instrument-playing scenes in the videos, it is …

Action Detection

Exploiting Fine-Grained Skip Behaviors for Micro-Video Recommendation

2025-04-04 · Sanghyuck Lee, Sangkeun Park, Jaesung Lee

The growing trend of sharing short videos on social media platforms, where users capture and share moments from their daily lives, has led to an increase in research efforts focused on micro-video recommendations. Howeve…

Micro-video recommendations