paper-with-me

홈 › Papers

Learning Fine-Grained Features for Pixel-wise Video Correspondences

2023-08-06 · ICCV 2023 1 · Rui Li, Shenglong Zhou, Dong Liu

Video analysis tasks rely heavily on identifying the pixels from different frames that correspond to the same visual target. To tackle this problem, recent studies have advocated feature learning methods that aim to learn distinctive representations to match the pixels, especially in a self-supervised fashion. Unfortunately, these methods have difficulties for tiny or even single-pixel visual targets. Pixel-wise video correspondences were traditionally related to optical flows, which however lead to deterministic correspondences and lack robustness on real-world videos. We address the problem of learning features for establishing pixel-wise correspondences. Motivated by optical flows as well as the self-supervised feature learning, we propose to use not only labeled synthetic videos but also unlabeled real-world videos for learning fine-grained representations in a holistic framework. We adopt an adversarial learning scheme to enhance the generalization ability of the learned features. Moreover, we design a coarse-to-fine framework to pursue high computational efficiency. Our experimental results on a series of correspondence-based tasks demonstrate that the proposed method outperforms state-of-the-art rivals in both accuracy and efficiency.

📄 PDF Abstract BibTeX arXiv:2308.03040

Code (1)

qianduoduolr/fgvc 공식 구현 pytorch

Tasks

Computational Efficiency

Similar Papers 제목 키워드 기반

Joint-task Self-supervised Learning for Temporal Correspondence

2019-09-26 · NeurIPS 2019 12 · Xueting Li, Sifei Liu, Shalini De Mello, Xiaolong Wang 외

This paper proposes to learn reliable dense correspondence from videos in a self-supervised manner. Our learning process integrates two highly related tasks: tracking large image regions \emph{and} establishing fine-grai…

Object TrackingSelf-Supervised LearningSemi-Supervised Video Object SegmentationUnsupervised Video Object Segmentation

Shallow Features Matter: Hierarchical Memory with Heterogeneous Interaction for Unsupervised Video Object Segmentation

2025-07-30 · Zheng Xiangyu, He Songcheng, Li Wanyun, Li Xiaoqiang 외 arxiv

Unsupervised Video Object Segmentation (UVOS) aims to predict pixel-level masks for the most salient objects in videos without any prior annotations. While memory mechanisms have been proven critical in various video seg…

Unsupervised Video Object SegmentationVideo Saliency DetectionVideo Segmentation

Motion-driven Visual Tempo Learning for Video-based Action Recognition

2022-02-24 · TIP 2022 5 · Yuanzhong Liu, Junsong Yuan, Zhigang Tu

Action visual tempo characterizes the dynamics and the temporal scale of an action, which is helpful to distinguish human actions that share high similarities in visual dynamics and appearance. Previous methods capture t…

Action Recognition

Toward Next-generation Medical Vision Backbones: Modeling Finer-grained Long-range Visual Dependency

2025-09-14 · Mingyuan Meng arxiv

Medical Image Computing (MIC) is a broad research topic covering both pixel-wise (e.g., segmentation, registration) and image-wise (e.g., classification, regression) vision tasks. Effective analysis demands models that c…

Long-range modeling

OmniAvatar: Efficient Audio-Driven Avatar Video Generation with Adaptive Body Animation

2025-06-23 · Qijun Gan, Ruizi Yang, Jianke Zhu, Shaofei Xue 외

Significant progress has been made in audio-driven human animation, while most existing methods focus mainly on facial movements, limiting their ability to create full-body animations with natural synchronization and flu…

Human AnimationVideo Generation