paper-with-me

홈 › Papers

NEVLP: Noise-Robust Framework for Efficient Vision-Language Pre-training

2024-09-15 · Yiyi Tao, Zhuoyue Wang, Hang Zhang, Lun Wang

The success of Vision Language Models (VLMs) on various vision-language tasks heavily relies on pre-training with large scale web-crawled datasets. However, the noisy and incomplete nature of web data makes dataset scale crucial for performance, rendering end-to-end training increasingly prohibitive. In this paper, we propose NEVLP, a noise-robust framework for efficient vision-language pre-training that requires less pre-training data. Specifically, we bridge the modality gap between a frozen image encoder and a large language model with a transformer and introduce two innovative learning strategies: noise-adaptive learning and concept-enhanced learning to mitigate the impact of noise. In noise-adaptive learning, we estimate the noise probability of each image-text pair based on the transformer's memorization effect and employ noise-adaptive regularization on image-text contrastive learning to condition cross-modal alignment. In concept-enhanced learning, we enrich incomplete text by incorporating visual concepts (objects in the image) to provide prior information about existing objects for image-text matching and image-grounded text generation, thereby mitigating text incompletion. Our framework effectively utilizes noisy web data and achieves state-of-the-art performance with less pre-training data across a wide range of vision-language tasks, including image-text retrieval, image captioning, and visual question answering.

📄 PDF Abstract BibTeX arXiv:2409.09582

Code (0)

등록된 구현이 없습니다.

Tasks

Contrastive Learningcross-modal alignmentImage CaptioningImage-text matchingImage-text RetrievalLanguage ModellingLarge Language ModelMemorizationQuestion AnsweringText GenerationText MatchingText RetrievalVisual Question Answering

Methods 이 논문이 사용한 방법론

Contrastive Learning 설명 없음

Similar Papers 제목 키워드 기반

JoAPR: Cleaning the Lens of Prompt Learning for Vision-Language Models

2024-01-01 · CVPR 2024 1 · Yuncheng Guo, Xiaodong Gu

Leveraging few-shot datasets in prompt learning for Vision-Language Models eliminates the need for manual prompt engineering while highlighting the necessity of accurate annotations for the labels. However high-level…

Prompt EngineeringPrompt Learning

SDAR-VL: Stable and Efficient Block-wise Diffusion for Vision-Language Understanding

2025-12-16 · Shuang Cheng, Yuhua Jiang, Zineng Zhou, Dawei Liu 외 arxiv

Block-wise discrete diffusion offers an attractive balance between parallel generation and causal dependency modeling, making it a promising backbone for vision-language modeling. However, its practical adoption has been…

Vision Large Language Models Are Good Noise Handlers in Engagement Analysis

2025-11-18 · Alexander Vedernikov, Puneet Kumar, Haoyu Chen, Tapio Seppänen 외 arxiv

Engagement recognition in video datasets, unlike traditional image classification tasks, is particularly challenged by subjective labels and noise limiting model performance. To overcome the challenges of subjective and …

Image Classification

Light-weight Fine-tuning Method for Defending Adversarial Noise in Pre-trained Medical Vision-Language Models

2024-07-02 · Xu Han, Linghao Jin, Xuezhe Ma, Xiaofeng Liu

Fine-tuning pre-trained Vision-Language Models (VLMs) has shown remarkable capabilities in medical image and textual depiction synergy. Nevertheless, many pre-training datasets are restricted by patient privacy concerns,…

Knowledge Boosting: Rethinking Medical Contrastive Vision-Language Pre-Training

2023-07-14 · Xiaofei Chen, Yuting He, Cheng Xue, Rongjun Ge 외

The foundation models based on pre-training technology have significantly advanced artificial intelligence from theoretical to practical applications. These models have facilitated the feasibility of computer-aided diagn…

Clinical KnowledgeDiagnosticRepresentation LearningRetrieval