paper-with-me

홈 › Papers

Violet: A Vision-Language Model for Arabic Image Captioning with Gemini Decoder

2023-11-15 · Abdelrahman Mohamed, Fakhraddin Alwajih, El Moatez Billah Nagoudi, Alcides Alcoba Inciarte, Muhammad Abdul-Mageed

Although image captioning has a vast array of applications, it has not reached its full potential in languages other than English. Arabic, for instance, although the native language of more than 400 million people, remains largely underrepresented in this area. This is due to the lack of labeled data and powerful Arabic generative models. We alleviate this issue by presenting a novel vision-language model dedicated to Arabic, dubbed \textit{Violet}. Our model is based on a vision encoder and a Gemini text decoder that maintains generation fluency while allowing fusion between the vision and language components. To train our model, we introduce a new method for automatically acquiring data from available English datasets. We also manually prepare a new dataset for evaluation. \textit{Violet} performs sizeably better than our baselines on all of our evaluation datasets. For example, it reaches a CIDEr score of $61.2$ on our manually annotated dataset and achieves an improvement of $13$ points on Flickr8k.

📄 PDF Abstract BibTeX arXiv:2311.08844

Code (0)

등록된 구현이 없습니다.

Tasks

DecoderImage CaptioningLanguage ModelingLanguage Modelling

Similar Papers 제목 키워드 기반

Bench-Marking And Improving Arabic Automatic Image Captioning Through The Use Of Multi-Task Learning Paradigm

2022-02-11 · Muhy Eddin Za'ter, Bashar Talafha

The continuous increase in the use of social media and the visual content on the internet have accelerated the research in computer vision field in general and the image captioning task in specific. The process of genera…

Image CaptioningMulti-Task LearningWord Embeddings

Multimodal Arabic Captioning with Interpretable Visual Concept Integration

2025-09-29 · Passant Elchafei, Amany Fashwan arxiv

We present VLCAP, an Arabic image captioning framework that integrates CLIP-based visual label retrieval with multimodal text generation. Rather than relying solely on end-to-end captioning, VLCAP grounds generation in i…

Image CaptioningText Generation

“Wikily” Supervised Neural Translation Tailored to Cross-Lingual Tasks

2021-11-01 · EMNLP 2021 11 · Mohammad Sadegh Rasooli, Chris Callison-Burch, Derry Tanti Wijaya

We present a simple but effective approach for leveraging Wikipedia for neural machine translation as well as cross-lingual tasks of image captioning and dependency parsing without using any direct supervision from exter…

Cross-Lingual TransferCross-Lingual Word EmbeddingsDependency ParsingImage Captioning+3

"Wikily" Supervised Neural Translation Tailored to Cross-Lingual Tasks

2021-04-16 · Mohammad Sadegh Rasooli, Chris Callison-Burch, Derry Tanti Wijaya

We present a simple but effective approach for leveraging Wikipedia for neural machine translation as well as cross-lingual tasks of image captioning and dependency parsing without using any direct supervision from exter…

Cross-Lingual TransferCross-Lingual Word EmbeddingsDependency ParsingImage Captioning+3

An Empirical Study of End-to-End Video-Language Transformers with Masked Visual Modeling

2022-09-04 · CVPR 2023 1 · Tsu-Jui Fu, Linjie Li, Zhe Gan, Kevin Lin 외

Masked visual modeling (MVM) has been recently proven effective for visual pre-training. While similar reconstructive objectives on video inputs (e.g., masked frame modeling) have been explored in video-language (VidL) p…

Fill MaskOptical Flow EstimationQuestion AnsweringRetrieval+8