paper-with-me

홈 › Papers

UniDiff: Advancing Vision-Language Models with Generative and Discriminative Learning

2023-06-01 · Xiao Dong, Runhui Huang, XiaoYong Wei, Zequn Jie, Jianxing Yu, Jian Yin, Xiaodan Liang

Recent advances in vision-language pre-training have enabled machines to perform better in multimodal object discrimination (e.g., image-text semantic alignment) and image synthesis (e.g., text-to-image generation). On the other hand, fine-tuning pre-trained models with discriminative or generative capabilities such as CLIP and Stable Diffusion on domain-specific datasets has shown to be effective in various tasks by adapting to specific domains. However, few studies have explored the possibility of learning both discriminative and generative capabilities and leveraging their synergistic effects to create a powerful and personalized multimodal model during fine-tuning. This paper presents UniDiff, a unified multi-modal model that integrates image-text contrastive learning (ITC), text-conditioned image synthesis learning (IS), and reciprocal semantic consistency modeling (RSC). UniDiff effectively learns aligned semantics and mitigates the issue of semantic collapse during fine-tuning on small datasets by leveraging RSC on visual features from CLIP and diffusion models, without altering the pre-trained model's basic architecture. UniDiff demonstrates versatility in both multi-modal understanding and generative tasks. Experimental results on three datasets (Fashion-man, Fashion-woman, and E-commercial Product) showcase substantial enhancements in vision-language retrieval and text-to-image generation, illustrating the advantages of combining discriminative and generative fine-tuning. The proposed UniDiff model establishes a robust pipeline for personalized modeling and serves as a benchmark for future comparisons in the field.

📄 PDF Abstract BibTeX arXiv:2306.00813

Code (0)

등록된 구현이 없습니다.

Tasks

Contrastive LearningImage GenerationRetrievalText to Image GenerationText-to-Image Generation

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…
CLIP Contrastive Language-Image Pre-training (CLIP), consisting of a simplified version of ConVIRT trained from scratch, is an efficient method of image representation learning…
Contrastive Learning 설명 없음

Similar Papers 제목 키워드 기반

One Transformer Fits All Distributions in Multi-Modal Diffusion at Scale

2023-03-12 · Fan Bao, Shen Nie, Kaiwen Xue, Chongxuan Li 외

This paper proposes a unified diffusion framework (dubbed UniDiffuser) to fit all distributions relevant to a set of multi-modal data in one model. Our key insight is -- learning diffusion models for marginal, conditiona…

AllImage GenerationImage to textText to Image Generation+1

Generative Language-Grounded Policy in Vision-and-Language Navigation with Bayes' Rule

2020-09-16 · ICLR 2021 1 · Shuhei Kurita, Kyunghyun Cho

Vision-and-language navigation (VLN) is a task in which an agent is embodied in a realistic 3D environment and follows an instruction to reach the goal node. While most of the previous studies have built and investigated…

Language ModelingLanguage ModellingVision and Language Navigation

UniDiff: Parameter-Efficient Adaptation of Diffusion Models for Land Cover Classification with Multi-Modal Remotely Sensed Imagery and Sparse Annotations

2025-11-29 · Yuzhen Hu, Saurabh Prasad arxiv

Sparse annotations fundamentally constrain multimodal remote sensing: even recent state-of-the-art supervised methods such as MSFMamba are limited by the availability of labeled data, restricting their practical deployme…

Sparse Attention Vectors: Generative Multimodal Model Features Are Discriminative Vision-Language Classifiers

2024-11-28 · Chancharik Mitra, Brandon Huang, Tianning Chai, Zhiqiu Lin 외

Generative Large Multimodal Models (LMMs) like LLaVA and Qwen-VL excel at a wide variety of vision-language (VL) tasks such as image captioning or visual question answering. Despite strong performance, LMMs are not direc…

Image Captioningimage-classificationImage ClassificationMultiple-choice+3

Phonology-Guided Speech-to-Speech Translation for African Languages

2024-10-30 · Peter Ochieng, Dennis Kaburu

We present a prosody-guided framework for speech-to-speech translation (S2ST) that aligns and translates speech \emph{without} transcripts by leveraging cross-linguistic pause synchrony. Analyzing a 6{,}000-hour East Afr…

Semantic SimilaritySemantic Textual SimilaritySpeech-to-Speech TranslationTranslation+1