paper-with-me

홈 › Papers

GTM: Gray Temporal Model for Video Recognition

2021-10-20 · Yanping Zhang, Yongxin Yu

Data input modality plays an important role in video action recognition. Normally, there are three types of input: RGB, flow stream and compressed data. In this paper, we proposed a new input modality: gray stream. Specifically, taken the stacked consecutive 3 gray images as input, which is the same size of RGB, can not only skip the conversion process from video decoding data to RGB, but also improve the spatio-temporal modeling ability at zero computation and zero parameters. Meanwhile, we proposed a 1D Identity Channel-wise Spatio-temporal Convolution(1D-ICSC) which captures the temporal relationship at channel-feature level within a controllable computation budget(by parameters G & R). Finally, we confirm its effectiveness and efficiency on several action recognition benchmarks, such as Kinetics, Something-Something, HMDB-51 and UCF-101, and achieve impressive results.

📄 PDF Abstract BibTeX arXiv:2110.10348

Code (0)

등록된 구현이 없습니다.

Tasks

Action RecognitionmodelTemporal Action LocalizationVideo Recognition

Similar Papers 제목 키워드 기반

Two Decades of Colorization and Decolorization for Images and Videos

2022-04-28 · Shiguang Liu

Colorization is a computer-aided process, which aims to give color to a gray image or video. It can be used to enhance black-and-white images, including black-and-white photos, old-fashioned films, and scientific imaging…

ColorizationImage EnhancementImage SegmentationSemantic Segmentation+1

Self-Supervised Face Presentation Attack Detection with Dynamic Grayscale Snippets

2022-08-27 · Usman Muhammad, Mourad Oussalah

Face presentation attack detection (PAD) plays an important role in defending face recognition systems against presentation attacks. The success of PAD largely relies on supervised learning that requires a huge number of…

Face Presentation Attack DetectionFace Recognitionmotion predictionRepresentation Learning

VideoPure: Diffusion-based Adversarial Purification for Video Recognition

2025-01-25 · Kaixun Jiang, Zhaoyu Chen, Jiyuan Fu, Lingyi Hong 외

Recent work indicates that video recognition models are vulnerable to adversarial examples, posing a serious security risk to downstream applications. However, current research has primarily focused on adversarial attack…

Adversarial DefenseAdversarial PurificationAdversarial RobustnessDenoising+1

Color When It Counts: Grayscale-Guided Online Triggering for Always-On Streaming Video Sensing

2026-03-23 · Weitong Cai, Hang Zhang, Yukai Huang, Shitong Sun 외 arxiv

Always-on sensing is essential for next-generation edge/wearable AI systems, yet continuous high-fidelity RGB video capture remains prohibitively expensive for resource-constrained mobile and edge platforms. We present a…

Learning to See Through a Baby's Eyes: Early Visual Diets Enable Robust Visual Intelligence in Humans and Machines

2025-11-18 · Yusen Cai, Qing Lin, Bhargava Satya Nunna, Mengmi Zhang arxiv

Newborns perceive the world with low-acuity, color-degraded, and temporally continuous vision, which gradually sharpens as infants develop. To explore the ecological advantages of such staged "visual diets", we train sel…

Self-Supervised LearningObject Recognition