Self-Supervised Learning of Motion Concepts by Optimizing Counterfactuals
Estimating motion in videos is an essential computer vision problem with many downstream applications, including controllable video generation and robotics. Current solutions are primarily trained using synthetic data or require tuning of situation-specific heuristics, which inherently limits these models' capabilities in real-world contexts. Despite recent developments in large-scale self-supervised learning from videos, leveraging such representations for motion estimation remains relatively underexplored. In this work, we develop Opt-CWM, a self-supervised technique for flow and occlusion estimation from a pre-trained next-frame prediction model. Opt-CWM works by learning to optimize counterfactual probes that extract motion information from a base video model, avoiding the need for fixed heuristics while training on unrestricted video inputs. We achieve state-of-the-art performance for motion estimation on real-world videos while requiring no labeled data.
Code (0)
등록된 구현이 없습니다.
Tasks
counterfactualMotion EstimationOcclusion EstimationSelf-Supervised LearningVideo GenerationMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Graph Edits for Counterfactual Explanations: A comparative study
Counterfactuals have been established as a popular explainability technique which leverages a set of minimal edits to alter the prediction of a classifier. When considering conceptual counterfactuals on images, the edits…
counterfactualGraph Neural NetworkKnowledge GraphsExplaining $\mathcal{ELH}$ Concept Descriptions through Counterfactual Reasoning
Knowledge bases are widely used for information management, enabling high-impact applications such as web search, question answering, and natural language processing. They also serve as the backbone for automatic decisio…
counterfactualCounterfactual ReasoningManagementQuestion AnsweringDeMoGen: Towards Decompositional Human Motion Generation with Energy-Based Diffusion Models
Human motions are compositional: complex behaviors can be described as combinations of simpler primitives. However, existing approaches primarily focus on forward modeling, e.g., learning holistic mappings from text to m…
Self-supervised Semantic Segmentation Grounded in Visual Concepts
Unsupervised semantic segmentation requires assigning a label to every pixel without any human annotations. Despite recent advances in self-supervised representation learning for individual images, unsupervised semantic …
Representation LearningSegmentationSelf-Supervised LearningSemantic Segmentation+1Towards generating more interpretable counterfactuals via concept vectors: a preliminary study on chest X-rays
An essential step in deploying medical imaging models is ensuring alignment with clinical knowledge and interpretability. We focus on mapping clinical concepts into the latent space of generative models to identify Conce…
Clinical Knowledge