paper-with-me

Papers

Exploring High-Order Self-Similarity for Video Understanding

2026-04-22 · Manjin Kim, Heeseung Kwon, Karteek Alahari, Minsu Cho arxiv

Space-time self-similarity (STSS), which captures visual correspondences across frames, provides an effective way to represent temporal dynamics for video understanding. In this work, we explore higher-order STSS and demonstrate how STSSs at different orders reveal distinct aspects of these dynamics. We then introduce the Multi-Order Self-Similarity (MOSS) module, a lightweight neural module designed to learn and integrate multi-order STSS features. It can be applied to diverse video tasks to enhance motion modeling capabilities while consuming only marginal computational cost and memory usage. Extensive experiments on video action recognition, motion-centric video VQA, and real-world robotic tasks consistently demonstrate substantial improvements, validating the broad applicability of MOSS as a general temporal modeling module. The source code and checkpoints will be publicly available.

📄 PDF Abstract BibTeX arXiv:2604.20760

Code (0)

등록된 구현이 없습니다.

Tasks

Action Recognition

Similar Papers 제목 키워드 기반

Tensor train rank minimization with nonlocal self-similarity for tensor completion

2020-04-29 · Meng Ding, Ting-Zhu Huang, Xi-Le Zhao, Michael K. Ng 외

The tensor train (TT) rank has received increasing attention in tensor completion due to its ability to capture the global correlation of high-order tensors ($\textrm{order} >3$). For third order visual data, direct TT r…

Exploring Global Diversity and Local Context for Video Summarization

2022-01-27 · Yingchao Pan, Ouhan Huang, Qinghao Ye, Zhongjin Li 외

Video summarization aims to automatically generate a diverse and concise summary which is useful in large-scale video processing. Most of the methods tend to adopt self-attention mechanism across video frames, which fail…

DiversityVideo Summarization

Exploring Temporal Granularity in Self-Supervised Video Representation Learning

2021-12-08 · Rui Qian, Yeqing Li, Liangzhe Yuan, Boqing Gong 외

This work presents a self-supervised learning framework named TeG to explore Temporal Granularity in learning video representations. In TeG, we sample a long clip from a video and a short clip that lies inside the long c…

Representation LearningSelf-Supervised Learning

Patch Spatio-Temporal Relation Prediction for Video Anomaly Detection

2024-03-28 · Hao Shen, Lu Shi, Wanru Xu, Yigang Cen 외

Video Anomaly Detection (VAD), aiming to identify abnormalities within a specific context and timeframe, is crucial for intelligent Video Surveillance Systems. While recent deep learning-based VAD models have shown promi…

Anomaly DetectionMulti-Label LearningPredictionRelation+3

Self-supervised Spatiotemporal Representation Learning by Exploiting Video Continuity

2021-12-11 · Hanwen Liang, Niamul Quader, Zhixiang Chi, Lizhe Chen 외

Recent self-supervised video representation learning methods have found significant success by exploring essential properties of videos, e.g. speed, temporal order, etc. This work exploits an essential yet under-explored…

Action LocalizationAction RecognitionRepresentation LearningRetrieval+1