paper-with-me

Papers

CONCAP: Seeing Beyond English with Concepts Retrieval-Augmented Captioning

2025-07-27 · George Ibrahim, Rita Ramos, Yova Kementchedjhieva arxiv

Multilingual vision-language models have made significant strides in image captioning, yet they still lag behind their English counterparts due to limited multilingual training data and costly large-scale model parameterization. Retrieval-augmented generation (RAG) offers a promising alternative by conditioning caption generation on retrieved examples in the target language, reducing the need for extensive multilingual training. However, multilingual RAG captioning models often depend on retrieved captions translated from English, which can introduce mismatches and linguistic biases relative to the source language. We introduce CONCAP, a multilingual image captioning model that integrates retrieved captions with image-specific concepts, enhancing the contextualization of the input image and grounding the captioning process across different languages. Experiments on the XM3600 dataset indicate that CONCAP enables strong performance on low- and mid-resource languages, with highly reduced data requirements. Our findings highlight the effectiveness of concept-aware retrieval augmentation in bridging multilingual performance gaps.

📄 PDF Abstract BibTeX arXiv:2507.20411

Code (0)

등록된 구현이 없습니다.

Tasks

Image Captioning

Similar Papers 제목 키워드 기반

The All-Seeing Project: Towards Panoptic Visual Recognition and Understanding of the Open World

2023-08-03 · Weiyun Wang, Min Shi, Qingyun Li, Wenhai Wang 외

We present the All-Seeing (AS) project: a large-scale data and model for recognizing and understanding everything in the open world. Using a scalable data engine that incorporates human feedback and efficient models in t…

AllQuestion AnsweringRetrievalText Retrieval

Measuring the Italian-English lexical gap for action verbs and its impact on translation

2017-04-01 · WS 2017 4 · Lorenzo Gregori, Aless Panunzi, ro

This paper describes a method to measure the lexical gap of action verbs in Italian and English by using the IMAGACT ontology of action. The fine-grained categorization of action concepts of the data source allowed to ha…

Translation

Med-CoReasoner: Reducing Language Disparities in Medical Reasoning via Language-Informed Co-Reasoning

2026-01-13 · Fan Gao, Sherry T. Tong, Jiwoong Sohn, Jiahao Huang 외 arxiv

While reasoning-enhanced large language models perform strongly on English medical tasks, a persistent multilingual gap remains, with substantially weaker reasoning in local languages, limiting equitable global medical d…

Natural Language InferenceClinical KnowledgeQuestion Answering

Cross-Lingual Transfer Learning for Speech Translation

2024-07-01 · Rao Ma, Mengjie Qian, Yassir Fathullah, Siyuan Tang 외

There has been increasing interest in building multilingual foundation models for NLP and speech research. This paper examines how to expand the speech translation capability of these models with restricted data. Whisper…

Cross-Lingual TransferDecoderspeech-recognitionSpeech Recognition+3

Seeing the Abstract: Translating the Abstract Language for Vision Language Models

2025-05-06 · CVPR 2025 1 · Davide Talon, Federico Girella, Ziyue Liu, Marco Cristani 외

Natural language goes beyond dryly describing visual content. It contains rich abstract concepts to express feeling, creativity and properties that cannot be directly perceived. Yet, current research in Vision Language M…

Image RetrievalRetrieval