Self-Supervised Flow Matching for Scalable Multi-Modal Synthesis
Strong semantic representations improve the convergence and generation quality of diffusion and flow models. Existing approaches largely rely on external models, which require separate training, operate on misaligned objectives, and exhibit unexpected scaling behavior. We argue that this dependence arises from the model's training objective, which poses a denoising task with little incentive to learn semantic representations. We introduce Self-Flow: a self-supervised flow matching paradigm that integrates representation learning within the generative framework. Our key mechanism, Dual-Timestep Scheduling, applies heterogeneous noise levels across tokens, creating an information asymmetry that forces the model to infer missing information from corrupted inputs. This drives learning strong representations alongside generative capabilities without external supervision. Our method generalizes across modalities and enables multi-modal training while following expected scaling laws, achieving superior image, video, and audio generation.
Code (0)
등록된 구현이 없습니다.
Tasks
Representation LearningAudio GenerationSimilar Papers 제목 키워드 기반
Self-Point-Flow: Self-Supervised Scene Flow Estimation from Point Clouds with Optimal Transport and Random Walk
Due to the scarcity of annotated scene flow data, self-supervised scene flow learning in point clouds has attracted increasing attention. In the self-supervised manner, establishing correspondences between two point clou…
Scene Flow EstimationSelf-Supervised LearningSelf-supervised Scene Flow EstimationFlow2Stereo: Effective Self-Supervised Learning of Optical Flow and Stereo Matching
In this paper, we propose a unified method to jointly learn optical flow and stereo matching. Our first intuition is stereo matching can be modeled as a special case of optical flow, and we can leverage 3D geometry behin…
3D geometryOptical Flow EstimationSelf-Supervised LearningStereo MatchingTilt Matching for Scalable Sampling and Fine-Tuning
We propose a simple, scalable algorithm for using stochastic interpolants to sample from unnormalized densities and for fine-tuning generative models. The approach, Tilt Matching, arises from a dynamical equation relatin…
Vector-Symbolic Architecture for Event-Based Optical Flow
From a perspective of feature matching, optical flow estimation for event cameras involves identifying event correspondences by comparing feature similarity across accompanying event frames. In this work, we introduces a…
Event-based Optical FlowOptical Flow EstimationSelf-Supervised LearningSelf-Assessed Generation: Trustworthy Label Generation for Optical Flow and Stereo Matching in Real-world
A significant challenge facing current optical flow and stereo methods is the difficulty in generalizing them well to the real world. This is mainly due to the high costs required to produce datasets, and the limitations…
Optical Flow EstimationStereo Matching