A Simple Data Augmentation Strategy for Text-in-Image Scientific VQA
Scientific visual question answering poses significant challenges for vision-language models due to the complexity of scientific figures and their multimodal context. Traditional approaches treat the figure and accompanying text (e.g., questions and answer options) as separate inputs. EXAMS-V introduced a new paradigm by embedding both visual and textual content into a single image. However, even state-of-the-art proprietary models perform poorly on this setup in zero-shot settings, underscoring the need for task-specific fine-tuning. To address the scarcity of training data in this "text-in-image" format, we synthesize a new dataset by converting existing separate image-text pairs into unified images. Fine-tuning a small multilingual multimodal model on a mix of our synthetic data and EXAMS-V yields notable gains across 13 languages, demonstrating strong average improvements and cross-lingual transfer.
Code (0)
등록된 구현이 없습니다.
Tasks
Visual Question AnsweringCross-Lingual TransferData AugmentationSimilar Papers 제목 키워드 기반
Cheap and Good? Simple and Effective Data Augmentation for Low Resource Machine Reading
We propose a simple and effective strategy for data augmentation for low-resource machine reading comprehension (MRC). Our approach first pretrains the answer extraction components of a MRC system on the augmented data t…
Data AugmentationMachine Reading ComprehensionReading ComprehensionRetrievalSegment Augmentation and Differentiable Ranking for Logo Retrieval
Logo retrieval is a challenging problem since the definition of similarity is more subjective compared to image retrieval tasks and the set of known similarities is very scarce. To tackle this challenge, in this paper, w…
Image RetrievalRetrievalSmart(Sampling)Augment: Optimal and Efficient Data Augmentation for Semantic Segmentation
Data augmentation methods enrich datasets with augmented data to improve the performance of neural networks. Recently, automated data augmentation methods have emerged, which automatically design augmentation strategies.…
Bayesian OptimizationData Augmentationimage-classificationImage Classification+5KeepAugment: A Simple Information-Preserving Data Augmentation Approach
Data augmentation (DA) is an essential technique for training state-of-the-art deep learning systems. In this paper, we empirically show data augmentation might introduce noisy augmented examples and consequently hurt th…
Data AugmentationGeneral Classificationimage-classificationImage Classification+3TextAug: Test time Text Augmentation for Multimodal Person Re-identification
Multimodal Person Reidentification is gaining popularity in the research community due to its effectiveness compared to counter-part unimodal frameworks. However, the bottleneck for multimodal deep learning is the need f…
Data AugmentationMultimodal Deep LearningPerson Re-IdentificationSentence+1