paper-with-me

홈 › Papers

VolDoGer: LLM-assisted Datasets for Domain Generalization in Vision-Language Tasks

2024-07-29 · Juhwan Choi, JuneHyoung Kwon, Jungmin Yun, Seunguk Yu, Youngbin Kim

Domain generalizability is a crucial aspect of a deep learning model since it determines the capability of the model to perform well on data from unseen domains. However, research on the domain generalizability of deep learning models for vision-language tasks remains limited, primarily because of the lack of required datasets. To address these challenges, we propose VolDoGer: Vision-Language Dataset for Domain Generalization, a dedicated dataset designed for domain generalization that addresses three vision-language tasks: image captioning, visual question answering, and visual entailment. We constructed VolDoGer by extending LLM-based data annotation techniques to vision-language tasks, thereby alleviating the burden of recruiting human annotators. We evaluated the domain generalizability of various models, ranging from fine-tuned models to a recent multimodal large language model, through VolDoGer.

📄 PDF Abstract BibTeX arXiv:2407.19795

Code (0)

등록된 구현이 없습니다.

Tasks

Deep LearningDomain GeneralizationImage CaptioningLanguage ModelingLanguage ModellingLarge Language ModelMultimodal Large Language ModelQuestion AnsweringVisual EntailmentVisual Question Answering

Similar Papers 제목 키워드 기반

WEDGE: Web-Image Assisted Domain Generalization for Semantic Segmentation

2021-09-29 · Namyup Kim, Taeyoung Son, Jaehyun Pahk, Cuiling Lan 외

Domain generalization for semantic segmentation is highly demanded in real applications, where a trained model is expected to work well in previously unseen domains. One challenge lies in the lack of data which could cov…

DiversityDomain GeneralizationSegmentationSemantic Segmentation

Co-Generation and Segmentation for Generalized Surgical Instrument Segmentation on Unlabelled Data

2021-03-16 · Megha Kalia, Tajwar Abrar Aleef, Nassir Navab, Septimiu E. Salcudean

Surgical instrument segmentation for robot-assisted surgery is needed for accurate instrument tracking and augmented reality overlays. Therefore, the topic has been the subject of a number of recent papers in the CAI com…

SegmentationTranslation

Resonant Anomaly Detection with Multiple Reference Datasets

2022-12-20 · Mayee F. Chen, Benjamin Nachman, Frederic Sala

An important class of techniques for resonant anomaly detection in high energy physics builds models that can distinguish between reference and target datasets, where only the latter has appreciable signal. Such techniqu…

Anomaly Detection

Generalized Radiograph Representation Learning via Cross-supervision between Images and Free-text Radiology Reports

2021-11-04 · Hong-Yu Zhou, Xiaoyu Chen, Yinghao Zhang, Ruibang Luo 외

Pre-training lays the foundation for recent successes in radiograph analysis supported by deep learning. It learns transferable image representations by conducting large-scale fully-supervised or self-supervised learning…

Representation LearningSelf-Supervised LearningTransfer Learning

LMT-GP: Combined Latent Mean-Teacher and Gaussian Process for Semi-supervised Low-light Image Enhancement

2024-08-29 · Ye Yu, Fengxin Chen, Jun Yu, Zhen Kan

While recent low-light image enhancement (LLIE) methods have made significant advancements, they still face challenges in terms of low visual quality and weak generalization ability when applied to complex scenarios. To …

GPRImage EnhancementLow-Light Image EnhancementPseudo Label