paper-with-me

홈 › Papers

Learning Self-Similarity in Space and Time as a Generalized Motion for Action Recognition

2021-01-01 · Heeseung Kwon, Manjin Kim, Suha Kwak, Minsu Cho

Spatio-temporal convolution often fails to learn motion dynamics in videos and thus an effective motion representation is required for video understanding in the wild. In this paper, we propose a rich and robust motion representation method based on spatio-temporal self-similarity (STSS). Given a sequence of frames, STSS represents each local region as similarities to its neighbors in space and time. By converting appearance features into relational values, it enables the learner to better recognize structural patterns in space and time. We leverage the whole volume of STSS and let our model learn to extract an effective motion representation from it. The proposed method is implemented as a neural block, dubbed SELFY, that can be easily inserted into neural architectures and learned end-to-end without additional supervision. With a sufficient volume of the neighborhood in space and time, it effectively captures long-term interaction and fast motion in the video, leading to robust action recognition. Our experimental analysis demonstrates its superiority over previous methods for motion modeling as well as its complementarity to spatio-temporal features from direct convolution. On the standard action recognition benchmarks, Something-Something-V1 & V2, Diving-48, and FineGym, the proposed method achieves the state-of-the-art results.

📄 PDF Abstract BibTeX

Code (1)

arunos728/SELFY 공식 구현 pytorch

Tasks

Action RecognitionVideo Understanding

Methods 이 논문이 사용한 방법론

Convolution A convolution is a type of matrix operation, consisting of a kernel, a small matrix of weights, that slides over input data performing element-wise multiplication with the…

Similar Papers 제목 키워드 기반

Learning Self-Similarity in Space and Time as Generalized Motion for Video Action Recognition

2021-02-14 · ICCV 2021 10 · Heeseung Kwon, Manjin Kim, Suha Kwak, Minsu Cho

Spatio-temporal convolution often fails to learn motion dynamics in videos and thus an effective motion representation is required for video understanding in the wild. In this paper, we propose a rich and robust motion r…

Action RecognitionTemporal Action LocalizationVideo Understanding

Exploring High-Order Self-Similarity for Video Understanding

2026-04-22 · Manjin Kim, Heeseung Kwon, Karteek Alahari, Minsu Cho arxiv

Space-time self-similarity (STSS), which captures visual correspondences across frames, provides an effective way to represent temporal dynamics for video understanding. In this work, we explore higher-order STSS and dem…

Action Recognition

Graph Multi-Similarity Learning for Molecular Property Prediction

2024-01-31 · Hao Xu, Zhengyang Zhou, Pengyu Hong

Enhancing accurate molecular property prediction relies on effective and proficient representation learning. It is crucial to incorporate diverse molecular relationships characterized by multi-similarity (self-similarity…

AttributeContrastive LearningDrug DiscoveryMolecular Property Prediction+4

Tempered Self-Similarity Alignment for Physically Plausible Video Generation

2026-05-24 · Manjin Kim, Suha Kwak, Minsu Cho arxiv

Despite remarkable advances in video generative models, they still struggle to generate physically realistic videos, frequently exhibiting appearance drift, implausible motion, and temporal inconsistencies. In this work,…

Video Generation

One-DoF Robotic Design of Overconstrained Limbs with Energy-Efficient, Self-Collision-Free Motion

2025-09-26 · Yuping Gu, Bangchao Huang, Haoran Sun, Ronghan Xu 외 arxiv

While it is expected to build robotic limbs with multiple degrees of freedom (DoF) inspired by nature, a single DoF design remains fundamental, providing benefits that include, but are not limited to, simplicity, robustn…