Busy-Quiet Video Disentangling for Video Classification
In video data, busy motion details from moving regions are conveyed within a specific frequency bandwidth in the frequency domain. Meanwhile, the rest of the frequencies of video data are encoded with quiet information with substantial redundancy, which causes low processing efficiency in existing video models that take as input raw RGB frames. In this paper, we consider allocating intenser computation for the processing of the important busy information and less computation for that of the quiet information. We design a trainable Motion Band-Pass Module (MBPM) for separating busy information from quiet information in raw video data. By embedding the MBPM into a two-pathway CNN architecture, we define a Busy-Quiet Net (BQN). The efficiency of BQN is determined by avoiding redundancy in the feature space processed by the two pathways: one operating on Quiet features of low-resolution, while the other processes Busy features. The proposed BQN outperforms many recent video processing models on Something-Something V1, Kinetics400, UCF101 and HMDB51 datasets.
Code (2)
Tasks
Action ClassificationAction RecognitionAction Recognition In VideosClassificationGeneral ClassificationVideo ClassificationSimilar Papers 제목 키워드 기반
BQN: Busy-Quiet Net Enabled by Motion Band-Pass Module for Action Recognition
A rich video data representation can be realized by means of spatio-temporal frequency analysis. In this research study we show that a video can be disentangled, following the learning of video characteristics accordin…
Action RecognitionBusyBot: Learning to Interact, Reason, and Plan in a BusyBoard Environment
We introduce BusyBoard, a toy-inspired robot learning environment that leverages a diverse set of articulated objects and inter-object functional relations to provide rich visual feedback for robot interactions. Based on…
Causal DiscoveryRobot ManipulationRobot Task PlanningScene Graph GenerationDisentangling Video with Independent Prediction
We propose an unsupervised variational model for disentangling video into independent factors, i.e. each factor's future can be predicted from its past without considering the others. We show that our approach often lear…
PredictionThe Un-Kidnappable Robot: Acoustic Localization of Sneaking People
How easy is it to sneak up on a robot? We examine whether we can detect people using only the incidental sounds they produce as they move, even when they try to be quiet. We collect a robotic dataset of high-quality 4-ch…
JADE: Joint Autoencoders for Dis-Entanglement
The problem of feature disentanglement has been explored in the literature, for the purpose of image and video processing and text analysis. State-of-the-art methods for disentangling feature representations rely on the …
DisentanglementGeneral Classification