paper-with-me

Papers

Learning Visual Generative Priors without Text

2024-12-10 · CVPR 2025 1 · Shuailei Ma, Kecheng Zheng, Ying WEI, Wei Wu, Fan Lu, Yifei Zhang, Chen-Wei Xie, Biao Gong, Jiapeng Zhu, Yujun Shen

Although text-to-image (T2I) models have recently thrived as visual generative priors, their reliance on high-quality text-image pairs makes scaling up expensive. We argue that grasping the cross-modality alignment is not a necessity for a sound visual generative prior, whose focus should be on texture modeling. Such a philosophy inspires us to study image-to-image (I2I) generation, where models can learn from in-the-wild images in a self-supervised manner. We first develop a pure vision-based training framework, Lumos, and confirm the feasibility and the scalability of learning I2I models. We then find that, as an upstream task of T2I, our I2I model serves as a more foundational visual prior and achieves on-par or better performance than existing T2I models using only 1/10 text-image pairs for fine-tuning. We further demonstrate the superiority of I2I priors over T2I priors on some text-irrelevant visual generative tasks, like image-to-3D and image-to-video. Our project page is available at https://xiaomabufei.github.io/lumos.

📄 PDF Abstract BibTeX arXiv:2412.07767

Code (0)

등록된 구현이 없습니다.

Tasks

Image to 3DPhilosophy

Methods 이 논문이 사용한 방법론

Focus 설명 없음

Similar Papers 제목 키워드 기반

Uni-AdaVD: Universal Concept Erasure for Visual Generation via Orthogonal Value Decomposition

2026-07-16 · Qifan Zhou, Yuan Wang, Yanbin Hao, Xiang Wang 외 arxiv

Visual generative models inevitably absorb undesirable concepts from uncurated pretraining data, making concept erasure essential for safe deployment. Existing erasure methods, however, are often architecture-specific an…

GenFirst: Generation Before Reconstruction for Stable End-to-End Latent Generative Modeling

2026-08-29 · Guangting Zheng, Yiyuan Zhang, Tao Yang, Yunpeng Chen 외 hf

Latent generative models typically follow a two-stage pipeline, training a variational autoencoder for reconstruction and then a generative model on the frozen latent space. Since reconstruction-optimized latents are not…

Text-to-Image GenerationRepresentation Learning

TAVAE: A VAE with Adaptable Priors Explains Contextual Modulation in the Visual Cortex

2026-02-12 · Balázs Meszéna, Keith T. Murray, Julien Corbo, O. Batuhan Erkat 외 arxiv

The brain interprets visual information through learned regularities, a computation formalized as probabilistic inference under a prior. The visual cortex establishes priors for this inference, some delivered through est…

Revisiting the Role of Language Priors in Vision-Language Models

2023-06-02 · Zhiqiu Lin, Xinyue Chen, Deepak Pathak, Pengchuan Zhang 외

Vision-language models (VLMs) are impactful in part because they can be applied to a variety of visual understanding tasks in a zero-shot fashion, without any fine-tuning. We study $\textit{generative VLMs}$ that are tra…

Image-text matchingImage-text RetrievalLanguage ModellingQuestion Answering+6

Composing diffusion priors with explicit physical context via generative Gibbs sampling

2026-05-11 · Weizhou Wang, Jonathan Weare, Aaron R. Dinner arxiv

Pretrained diffusion models provide powerful learned priors, but in scientific sampling the target distribution often depends on physical context that is not fully represented by one generative model. We introduce Genera…