paper-with-me

홈 › Papers

Unifying Vision-Language Latents for Zero-label Image Caption Enhancement

2025-10-14 · Sanghyun Byun, Jung Ick Guack, Mohanad Odema, Baisub Lee, Jacob Song, Woo Seong Chung arxiv

Vision-language models (VLMs) achieve remarkable performance through large-scale image-text pretraining. However, their reliance on labeled image datasets limits scalability and leaves vast amounts of unlabeled image data underutilized. To address this, we propose Unified Vision-Language Alignment for Zero-Label Enhancement (ViZer), an enhancement training framework that enables zero-label learning in image captioning, providing a practical starting point for broader zero-label adaptation in vision-language tasks. Unlike prior approaches that rely on human or synthetically annotated datasets, ViZer actively aligns vision and language representation features during training, enabling existing VLMs to generate improved captions without requiring text labels or full retraining. We demonstrate ViZer's advantage in qualitative evaluation, as automated caption metrics such as CIDEr and BERTScore often penalize details that are absent in reference captions. Applying ViZer on SmolVLM-Base and Qwen2-VL, we observe consistent qualitative improvements, producing captions that are more grounded and descriptive than their baseline.

📄 PDF Abstract BibTeX arXiv:2510.12931

Code (0)

등록된 구현이 없습니다.

Tasks

Image Captioning

Similar Papers 제목 키워드 기반

Semantically Grounded QFormer for Efficient Vision Language Understanding

2023-11-13 · Moulik Choraria, Xinbo Wu, Sourya Basu, Nitesh Sekhar 외

General purpose Vision Language Models (VLMs) have received tremendous interest in recent years, owing to their ability to learn rich vision-language correlations as well as their broad zero-shot competencies. One immens…

DiversityImage to textRepresentation LearningText Generation

Capturing Label Characteristics in VAEs

2020-06-17 · ICLR 2021 1 · Tom Joy, Sebastian M. Schmon, Philip H. S. Torr, N. Siddharth 외

We present a principled approach to incorporating labels in VAEs that captures the rich characteristic information associated with those labels. While prior work has typically conflated these by learning latent variables…

Enhanced Continual Learning of Vision-Language Models with Model Fusion

2025-03-12 · Haoyuan Gao, Zicong Zhang, Yuqi Wei, Linglan Zhao 외

Vision-Language Models (VLMs) represent a breakthrough in artificial intelligence by integrating visual and textual modalities to achieve impressive zero-shot capabilities. However, VLMs are susceptible to catastrophic f…

Continual Learningparameter-efficient fine-tuning

Unifying (Machine) Vision via Counterfactual World Modeling

2023-06-02 · Daniel M. Bear, Kevin Feigelis, Honglin Chen, Wanhee Lee 외

Leading approaches in machine vision employ different architectures for different tasks, trained on costly task-specific labeled datasets. This complexity has held back progress in areas, such as robotics, where robust t…

counterfactualOptical Flow Estimation

GeoLangBind: Unifying Earth Observation with Agglomerative Vision-Language Foundation Models

2025-03-08 · Zhitong Xiong, Yi Wang, Weikang Yu, Adam J Stewart 외

Earth observation (EO) data, collected from diverse sensors with varying imaging principles, present significant challenges in creating unified analytical frameworks. We present GeoLangBind, a novel agglomerative vision-…

Earth Observation