paper-with-me

Papers Video Action Detection

“Video Action Detection” 태그가 달린 논문 32편 · 필터 해제

Scaling Open-Vocabulary Action Detection

2025-04-04 · Zhen Hao Sia, Yogesh Singh Rawat

In this work, we focus on scaling open-vocabulary action detection. Existing approaches for action detection are predominantly limited to closed-set scenarios and rely on complex, parameter-heavy architectures. Extending…

Action DetectionMultiple Action DetectionOpen Vocabulary Action DetectionSpatio-Temporal Action Localization+2

JoVALE: Detecting Human Actions in Video Using Audiovisual and Language Contexts

2024-12-18 · Taein Son, Soo Won Seo, Jisong Kim, Seok Hwan Lee 외

Video Action Detection (VAD) entails localizing and categorizing action instances within videos, which inherently consist of diverse information sources such as audio, visual cues, and surrounding scene contexts. Leverag…

Action DetectionDescriptiveImage CaptioningVideo Action Detection

Stable Mean Teacher for Semi-supervised Video Action Detection

2024-12-10 · Akash Kumar, Sirshapan Mitra, Yogesh Singh Rawat

In this work, we focus on semi-supervised learning for video action detection. Video action detection requires spatiotemporal localization in addition to classification, and a limited amount of labels makes the model pro…

Action DetectionSemantic SegmentationSemi_supervised Video Action DetectionSemi-Supervised Video Action Detection+3

On Occlusions in Video Action Detection: Benchmark Datasets And Training Recipes

2024-10-25 · NeurIPS 2023 11 · Rajat Modi, Vibhav Vineet, Yogesh Singh Rawat

This paper explores the impact of occlusions in video action detection. We facilitate this study by introducing five new benchmark datasets namely O-UCF and O-JHMDB consisting of synthetically controlled static/dynamic o…

Action DetectionData AugmentationVideo Action Detection

JARViS: Detecting Actions in Video Using Unified Actor-Scene Context Relation Modeling

2024-08-07 · Seok Hwan Lee, Taein Son, Soo Won Seo, Jisong Kim 외

Video action detection (VAD) is a formidable vision task that involves the localization and classification of actions within the spatial and temporal dimensions of a video clip. Among the myriad VAD architectures, two-st…

Action DetectionRelationVideo Action Detection

Classification Matters: Improving Video Action Detection with Class-Specific Attention

2024-07-29 · Jinsung Lee, Taeoh Kim, Inwoong Lee, Minho Shim 외

Video action detection (VAD) aims to detect actors and classify their actions in a video. We figure that VAD suffers more from classification rather than localization of actors. Hence, we analyze how prevailing methods f…

Action DetectionClassificationVideo Action Detection

Benchmarking Deep Learning Models on NVIDIA Jetson Nano for Real-Time Systems: An Empirical Investigation

2024-06-25 · Tushar Prasanna Swaminathan, Christopher Silver, Thangarajah Akilan

The proliferation of complex deep learning (DL) models has revolutionized various applications, including computer vision-based solutions, prompting their integration into real-time systems. However, the resource-intensi…

Action DetectionBenchmarkingimage-classificationImage Classification+2

Generative Model-based Feature Knowledge Distillation for Action Recognition

2023-12-14 · Guiqin Wang, Peng Zhao, Yanjiang Shi, Cong Zhao 외

Knowledge distillation (KD), a technique widely employed in computer vision, has emerged as a de facto standard for improving the performance of small neural networks. However, prevailing KD-based approaches in video tas…

Action DetectionAction RecognitionKnowledge DistillationModel Compression+2

Semi-supervised Active Learning for Video Action Detection

2023-12-12 · Ayush Singh, Aayush J Rana, Akash Kumar, Shruti Vyas 외

In this work, we focus on label efficient learning for video action detection. We develop a novel semi-supervised active learning approach which utilizes both labeled as well as unlabeled data along with informative samp…

Action DetectionActive LearningPseudo LabelSemantic Segmentation+5

A Grammatical Compositional Model for Video Action Detection

2023-10-04 · Zhijun Zhang, Xu Zou, Jiahuan Zhou, Sheng Zhong 외

Analysis of human actions in videos demands understanding complex human dynamics, as well as the interaction between actors and context. However, these interaction relationships usually exhibit large intra-class variatio…

Action DetectionHuman DynamicsmodelVideo Action Detection

M$^{3}$3D: Learning 3D priors using Multi-Modal Masked Autoencoders for 2D image and video understanding

2023-09-26 · Muhammad Abdullah Jamal, Omid Mohareri

We present a new pre-training strategy called M$^{3}$3D ($\underline{M}$ulti-$\underline{M}$odal $\underline{M}$asked $\underline{3D}$) built based on Multi-modal masked autoencoders that can leverage 3D priors and learn…

2D Semantic SegmentationAction DetectionAction RecognitionContrastive Learning+7

MRSN: Multi-Relation Support Network for Video Action Detection

2023-04-24 · Yin-Dong Zheng, Guo Chen, Minglei Yuan, Tong Lu

Action detection is a challenging video understanding task, requiring modeling spatio-temporal and interaction relations. Current methods usually model actor-actor and actor-context relations separately, ignoring their c…

Action DetectionRelationVideo Action DetectionVideo Understanding

Efficient Video Action Detection with Token Dropout and Context Refinement

2023-04-17 · ICCV 2023 1 · Lei Chen, Zhan Tong, Yibing Song, Gangshan Wu 외

Streaming video clips with large-scale video tokens impede vision transformers (ViTs) for efficient recognition, especially in video action detection where sufficient spatiotemporal representations are required for preci…

Action DetectionDecoderVideo Action Detection

CycleACR: Cycle Modeling of Actor-Context Relations for Video Action Detection

2023-03-28 · Lei Chen, Zhan Tong, Yibing Song, Gangshan Wu 외

The relation modeling between actors and scene context advances video action detection where the correlation of multiple actors makes their action recognition challenging. Existing studies model each actor and scene rela…

Action DetectionAction RecognitionRelationRelation Network+1

Context Understanding in Computer Vision: A Survey

2023-02-10 · Xuan Wang, Zhigang Zhu

Contextual information plays an important role in many computer vision tasks, such as object detection, video action detection, image classification, etc. Recognizing a single object or action out of context could be som…

Action Detectionimage-classificationImage ClassificationIn-Context Learning+5

Hybrid Active Learning via Deep Clustering for Video Action Detection

2023-01-01 · CVPR 2023 1 · Aayush J. Rana, Yogesh S. Rawat

In this work, we focus on reducing the annotation cost for video action detection which requires costly frame-wise dense annotations. We study a novel hybrid active learning (AL) strategy which performs efficient lab…

Action DetectionActive LearningClusteringDeep Clustering+3

Spotting Temporally Precise, Fine-Grained Events in Video

2022-07-20 · James Hong, Haotian Zhang, Michaël Gharbi, Matthew Fisher 외

We introduce the task of spotting temporally precise, fine-grained events in video (detecting the precise moment in time events occur). Precise spotting requires models to reason globally about the full-time scale of act…

Action DetectionAction SpottingGPUSegmentation+2

Video Action Detection: Analysing Limitations and Challenges

2022-04-17 · Rajat Modi, Aayush Jung Rana, Akash Kumar, Praveen Tirupattur 외

Beyond possessing large enough size to feed data hungry machines (eg, transformers), what attributes measure the quality of a dataset? Assuming that the definitions of such attributes do exist, how do we quantify among t…

Action DetectionVideo Action Detection

E^2TAD: An Energy-Efficient Tracking-based Action Detector

2022-04-09 · Xin Hu, Zhenyu Wu, Hao-Yu Miao, Siqi Fan 외

Video action detection (spatio-temporal action localization) is usually the starting point for human-centric intelligent analysis of videos nowadays. It has high practical impacts for many applications across robotics, s…

Action DetectionAction LocalizationFine-Grained Action Detectionobject-detection+4

Context-LSTM: a robust classifier for video detection on UCF101

2022-03-13 · Dengshan Li, Rujing Wang

Video detection and human action recognition may be computationally expensive, and need a long time to train models. In this paper, we were intended to reduce the training time and the GPU memory usage of video detection…

Action DetectionAction RecognitionGPUTemporal Action Localization+1
1–20 / 32 다음 →