Papers Video Action Detection
“Video Action Detection” 태그가 달린 논문 32편 · 필터 해제
Scaling Open-Vocabulary Action Detection
In this work, we focus on scaling open-vocabulary action detection. Existing approaches for action detection are predominantly limited to closed-set scenarios and rely on complex, parameter-heavy architectures. Extending…
Action DetectionMultiple Action DetectionOpen Vocabulary Action DetectionSpatio-Temporal Action Localization+2JoVALE: Detecting Human Actions in Video Using Audiovisual and Language Contexts
Video Action Detection (VAD) entails localizing and categorizing action instances within videos, which inherently consist of diverse information sources such as audio, visual cues, and surrounding scene contexts. Leverag…
Action DetectionDescriptiveImage CaptioningVideo Action DetectionStable Mean Teacher for Semi-supervised Video Action Detection
In this work, we focus on semi-supervised learning for video action detection. Video action detection requires spatiotemporal localization in addition to classification, and a limited amount of labels makes the model pro…
Action DetectionSemantic SegmentationSemi_supervised Video Action DetectionSemi-Supervised Video Action Detection+3On Occlusions in Video Action Detection: Benchmark Datasets And Training Recipes
This paper explores the impact of occlusions in video action detection. We facilitate this study by introducing five new benchmark datasets namely O-UCF and O-JHMDB consisting of synthetically controlled static/dynamic o…
Action DetectionData AugmentationVideo Action DetectionJARViS: Detecting Actions in Video Using Unified Actor-Scene Context Relation Modeling
Video action detection (VAD) is a formidable vision task that involves the localization and classification of actions within the spatial and temporal dimensions of a video clip. Among the myriad VAD architectures, two-st…
Action DetectionRelationVideo Action DetectionClassification Matters: Improving Video Action Detection with Class-Specific Attention
Video action detection (VAD) aims to detect actors and classify their actions in a video. We figure that VAD suffers more from classification rather than localization of actors. Hence, we analyze how prevailing methods f…
Action DetectionClassificationVideo Action DetectionBenchmarking Deep Learning Models on NVIDIA Jetson Nano for Real-Time Systems: An Empirical Investigation
The proliferation of complex deep learning (DL) models has revolutionized various applications, including computer vision-based solutions, prompting their integration into real-time systems. However, the resource-intensi…
Action DetectionBenchmarkingimage-classificationImage Classification+2Generative Model-based Feature Knowledge Distillation for Action Recognition
Knowledge distillation (KD), a technique widely employed in computer vision, has emerged as a de facto standard for improving the performance of small neural networks. However, prevailing KD-based approaches in video tas…
Action DetectionAction RecognitionKnowledge DistillationModel Compression+2Semi-supervised Active Learning for Video Action Detection
In this work, we focus on label efficient learning for video action detection. We develop a novel semi-supervised active learning approach which utilizes both labeled as well as unlabeled data along with informative samp…
Action DetectionActive LearningPseudo LabelSemantic Segmentation+5A Grammatical Compositional Model for Video Action Detection
Analysis of human actions in videos demands understanding complex human dynamics, as well as the interaction between actors and context. However, these interaction relationships usually exhibit large intra-class variatio…
Action DetectionHuman DynamicsmodelVideo Action DetectionM$^{3}$3D: Learning 3D priors using Multi-Modal Masked Autoencoders for 2D image and video understanding
We present a new pre-training strategy called M$^{3}$3D ($\underline{M}$ulti-$\underline{M}$odal $\underline{M}$asked $\underline{3D}$) built based on Multi-modal masked autoencoders that can leverage 3D priors and learn…
2D Semantic SegmentationAction DetectionAction RecognitionContrastive Learning+7MRSN: Multi-Relation Support Network for Video Action Detection
Action detection is a challenging video understanding task, requiring modeling spatio-temporal and interaction relations. Current methods usually model actor-actor and actor-context relations separately, ignoring their c…
Action DetectionRelationVideo Action DetectionVideo UnderstandingEfficient Video Action Detection with Token Dropout and Context Refinement
Streaming video clips with large-scale video tokens impede vision transformers (ViTs) for efficient recognition, especially in video action detection where sufficient spatiotemporal representations are required for preci…
Action DetectionDecoderVideo Action DetectionCycleACR: Cycle Modeling of Actor-Context Relations for Video Action Detection
The relation modeling between actors and scene context advances video action detection where the correlation of multiple actors makes their action recognition challenging. Existing studies model each actor and scene rela…
Action DetectionAction RecognitionRelationRelation Network+1Context Understanding in Computer Vision: A Survey
Contextual information plays an important role in many computer vision tasks, such as object detection, video action detection, image classification, etc. Recognizing a single object or action out of context could be som…
Action Detectionimage-classificationImage ClassificationIn-Context Learning+5Hybrid Active Learning via Deep Clustering for Video Action Detection
In this work, we focus on reducing the annotation cost for video action detection which requires costly frame-wise dense annotations. We study a novel hybrid active learning (AL) strategy which performs efficient lab…
Action DetectionActive LearningClusteringDeep Clustering+3Spotting Temporally Precise, Fine-Grained Events in Video
We introduce the task of spotting temporally precise, fine-grained events in video (detecting the precise moment in time events occur). Precise spotting requires models to reason globally about the full-time scale of act…
Action DetectionAction SpottingGPUSegmentation+2Video Action Detection: Analysing Limitations and Challenges
Beyond possessing large enough size to feed data hungry machines (eg, transformers), what attributes measure the quality of a dataset? Assuming that the definitions of such attributes do exist, how do we quantify among t…
Action DetectionVideo Action DetectionE^2TAD: An Energy-Efficient Tracking-based Action Detector
Video action detection (spatio-temporal action localization) is usually the starting point for human-centric intelligent analysis of videos nowadays. It has high practical impacts for many applications across robotics, s…
Action DetectionAction LocalizationFine-Grained Action Detectionobject-detection+4Context-LSTM: a robust classifier for video detection on UCF101
Video detection and human action recognition may be computationally expensive, and need a long time to train models. In this paper, we were intended to reduce the training time and the GPU memory usage of video detection…
Action DetectionAction RecognitionGPUTemporal Action Localization+1