Violet: A Vision-Language Model for Arabic Image Captioning with Gemini Decoder
Although image captioning has a vast array of applications, it has not reached its full potential in languages other than English. Arabic, for instance, although the native language of more than 400 million people, remains largely underrepresented in this area. This is due to the lack of labeled data and powerful Arabic generative models. We alleviate this issue by presenting a novel vision-language model dedicated to Arabic, dubbed \textit{Violet}. Our model is based on a vision encoder and a Gemini text decoder that maintains generation fluency while allowing fusion between the vision and language components. To train our model, we introduce a new method for automatically acquiring data from available English datasets. We also manually prepare a new dataset for evaluation. \textit{Violet} performs sizeably better than our baselines on all of our evaluation datasets. For example, it reaches a CIDEr score of $61.2$ on our manually annotated dataset and achieves an improvement of $13$ points on Flickr8k.
Code (0)
등록된 구현이 없습니다.
Tasks
DecoderImage CaptioningLanguage ModelingLanguage ModellingSimilar Papers 제목 키워드 기반
Bench-Marking And Improving Arabic Automatic Image Captioning Through The Use Of Multi-Task Learning Paradigm
The continuous increase in the use of social media and the visual content on the internet have accelerated the research in computer vision field in general and the image captioning task in specific. The process of genera…
Image CaptioningMulti-Task LearningWord EmbeddingsMultimodal Arabic Captioning with Interpretable Visual Concept Integration
We present VLCAP, an Arabic image captioning framework that integrates CLIP-based visual label retrieval with multimodal text generation. Rather than relying solely on end-to-end captioning, VLCAP grounds generation in i…
Image CaptioningText Generation“Wikily” Supervised Neural Translation Tailored to Cross-Lingual Tasks
We present a simple but effective approach for leveraging Wikipedia for neural machine translation as well as cross-lingual tasks of image captioning and dependency parsing without using any direct supervision from exter…
Cross-Lingual TransferCross-Lingual Word EmbeddingsDependency ParsingImage Captioning+3"Wikily" Supervised Neural Translation Tailored to Cross-Lingual Tasks
We present a simple but effective approach for leveraging Wikipedia for neural machine translation as well as cross-lingual tasks of image captioning and dependency parsing without using any direct supervision from exter…
Cross-Lingual TransferCross-Lingual Word EmbeddingsDependency ParsingImage Captioning+3An Empirical Study of End-to-End Video-Language Transformers with Masked Visual Modeling
Masked visual modeling (MVM) has been recently proven effective for visual pre-training. While similar reconstructive objectives on video inputs (e.g., masked frame modeling) have been explored in video-language (VidL) p…
Fill MaskOptical Flow EstimationQuestion AnsweringRetrieval+8