paper-with-me

홈 › Papers

Exploring Temporally Dynamic Data Augmentation for Video Recognition

2022-06-30 · Taeoh Kim, Jinhyung Kim, Minho Shim, Sangdoo Yun, Myunggu Kang, Dongyoon Wee, Sangyoun Lee

Data augmentation has recently emerged as an essential component of modern training recipes for visual recognition tasks. However, data augmentation for video recognition has been rarely explored despite its effectiveness. Few existing augmentation recipes for video recognition naively extend the image augmentation methods by applying the same operations to the whole video frames. Our main idea is that the magnitude of augmentation operations for each frame needs to be changed over time to capture the real-world video's temporal variations. These variations should be generated as diverse as possible using fewer additional hyper-parameters during training. Through this motivation, we propose a simple yet effective video data augmentation framework, DynaAugment. The magnitude of augmentation operations on each frame is changed by an effective mechanism, Fourier Sampling that parameterizes diverse, smooth, and realistic temporal variations. DynaAugment also includes an extended search space suitable for video for automatic data augmentation methods. DynaAugment experimentally demonstrates that there are additional performance rooms to be improved from static augmentations on diverse video models. Specifically, we show the effectiveness of DynaAugment on various video datasets and tasks: large-scale video recognition (Kinetics-400 and Something-Something-v2), small-scale video recognition (UCF- 101 and HMDB-51), fine-grained video recognition (Diving-48 and FineGym), video action segmentation on Breakfast, video action localization on THUMOS'14, and video object detection on MOT17Det. DynaAugment also enables video models to learn more generalized representation to improve the model robustness on the corrupted videos.

📄 PDF Abstract BibTeX arXiv:2206.15015

Code (0)

등록된 구현이 없습니다.

Tasks

Action LocalizationAction SegmentationData AugmentationImage Augmentationobject-detectionObject DetectionVideo Object DetectionVideo Recognition

Similar Papers 제목 키워드 기반

Learning Temporally Invariant and Localizable Features via Data Augmentation for Video Recognition

2020-08-13 · Taeoh Kim, Hyeongmin Lee, MyeongAh Cho, Ho Seong Lee 외

Deep-Learning-based video recognition has shown promising improvements along with the development of large-scale datasets and spatiotemporal network architectures. In image recognition, learning spatially invariant featu…

Action RecognitionData AugmentationVideo Recognition

Single-frame Regularization for Temporally Stable CNNs

2019-02-27 · CVPR 2019 6 · Gabriel Eilertsen, Rafał K. Mantiuk, Jonas Unger

Convolutional neural networks (CNNs) can model complicated non-linear relations between images. However, they are notoriously sensitive to small changes in the input. Most CNNs trained to describe image-to-image mappings…

Motion EstimationOptical Flow Estimation

BodyReLux: Temporally Consistent Full-Body Video Relighting

2026-05-20 · Li Ma, Mingming He, Xueming Yu, David M. George 외 arxiv

Being able to relight human performance is a fundamental task for post production and content creation. We present BodyReLux, a subject-specific video diffusion-based framework for relighting full-body human performances…

Data Augmentation

Temporally Consistent Dynamic Scene Graphs: An End-to-End Approach for Action Tracklet Generation

2024-12-03 · Raphael Ruschel, Md Awsafur Rahman, Hardik Prajapati, Suya You 외

Understanding video content is pivotal for advancing real-world applications like activity recognition, autonomous systems, and human-computer interaction. While scene graphs are adept at capturing spatial relationships …

Activity RecognitionAutonomous NavigationDecoder

Vivid4D: Improving 4D Reconstruction from Monocular Video by Video Inpainting

2025-04-15 · Jiaxin Huang, Sheng Miao, Bangbang Yang, Yuewen Ma 외

Reconstructing 4D dynamic scenes from casually captured monocular videos is valuable but highly challenging, as each timestamp is observed from a single viewpoint. We introduce Vivid4D, a novel approach that enhances 4D …

4D reconstructionVideo Inpainting