paper-with-me

홈 › Papers

Fully-Coupled Two-Stream Spatiotemporal Networks for Extremely Low Resolution Action Recognition

2018-01-11 · Mingze Xu, Aidean Sharghi, Xin Chen, David J. Crandall

A major emerging challenge is how to protect people's privacy as cameras and computer vision are increasingly integrated into our daily lives, including in smart devices inside homes. A potential solution is to capture and record just the minimum amount of information needed to perform a task of interest. In this paper, we propose a fully-coupled two-stream spatiotemporal architecture for reliable human action recognition on extremely low resolution (e.g., 12x16 pixel) videos. We provide an efficient method to extract spatial and temporal features and to aggregate them into a robust feature representation for an entire action video sequence. We also consider how to incorporate high resolution videos during training in order to build better low resolution action recognition models. We evaluate on two publicly-available datasets, showing significant improvements over the state-of-the-art.

📄 PDF Abstract BibTeX arXiv:1801.03983

Code (0)

등록된 구현이 없습니다.

Tasks

Action RecognitionTemporal Action Localization

Similar Papers 제목 키워드 기반

Semi-Coupled Two-Stream Fusion ConvNets for Action Recognition at Extremely Low Resolutions

2016-10-12 · Jiawei Chen, Jonathan Wu, Janusz Konrad, Prakash Ishwar

Deep convolutional neural networks (ConvNets) have been recently shown to attain state-of-the-art performance for action recognition on standard-resolution videos. However, less attention has been paid to recognition per…

Action RecognitionTemporal Action Localization

MATEY: multiscale adaptive foundation models for spatiotemporal physical systems

2024-12-29 · Pei Zhang, M. Paul Laiu, Matthew Norman, Doug Stefanski 외

Accurate representation of the multiscale features in spatiotemporal physical systems using vision transformer (ViT) architectures requires extremely long, computationally prohibitive token sequences. To address this iss…

Computational Efficiency

Ultrafast Video Attention Prediction with Coupled Knowledge Distillation

2019-04-09 · Kui Fu, Peipei Shi, Yafei Song, Shiming Ge 외

Large convolutional neural network models have recently demonstrated impressive performance on video attention prediction. Conventionally, these models are with intensive computation and large memory. To address these is…

CPUGPUKnowledge DistillationPrediction

Self-Supervised Video Representation Learning with Constrained Spatiotemporal Jigsaw

2021-01-01 · Yuqi Huo, Mingyu Ding, Haoyu Lu, Zhiwu Lu 외

This paper proposes a novel pretext task for self-supervised video representation learning by exploiting spatiotemporal continuity in videos. It is motivated by the fact that videos are spatiotemporal by nature and a rep…

Representation Learning

Ultra Flash: Scaling Real-Time Streaming Video Generation to High Resolutions

2026-06-08 · Luxury, Jie Huang, Zihao Fan, Xiaoxiao Ma 외 arxiv

While recent autoregressive video diffusion models achieve remarkable streaming quality, they remain confined to low resolutions (e.g., 480P), leaving efficient, scalable, real-time high-resolution video generation a fun…

Video Generation