paper-with-me

홈 › Papers

Improving Captioning for Low-Resource Languages by Cycle Consistency

2019-08-21 · Yike Wu, Shiwan Zhao, Jia Chen, Ying Zhang, Xiaojie Yuan, Zhong Su

Improving the captioning performance on low-resource languages by leveraging English caption datasets has received increasing research interest in recent years. Existing works mainly fall into two categories: translation-based and alignment-based approaches. In this paper, we propose to combine the merits of both approaches in one unified architecture. Specifically, we use a pre-trained English caption model to generate high-quality English captions, and then take both the image and generated English captions to generate low-resource language captions. We improve the captioning performance by adding the cycle consistency constraint on the cycle of image regions, English words, and low-resource language words. Moreover, our architecture has a flexible design which enables it to benefit from large monolingual English caption datasets. Experimental results demonstrate that our approach outperforms the state-of-the-art methods on common evaluation metrics. The attention visualization also shows that the proposed approach really improves the fine-grained alignment between words and image regions.

📄 PDF Abstract BibTeX arXiv:1908.07810

Code (0)

등록된 구현이 없습니다.

Tasks

Translation

Similar Papers 제목 키워드 기반

CycleCap: Improving VLMs Captioning Performance via Self-Supervised Cycle Consistency Fine-Tuning

2026-03-18 · Marios Krestenitis, Christos Tzelepis, Konstantinos Ioannidis, Stefanos Vrochidis 외 arxiv

Visual-Language Models (VLMs) have achieved remarkable progress in image captioning, visual question answering, and visual reasoning. Yet they remain prone to vision-language misalignment, often producing overly generic …

Visual Question AnsweringVisual ReasoningImage Captioning

Viewpoint-Agnostic Change Captioning With Cycle Consistency

2021-01-01 · ICCV 2021 10 · Hoeseong Kim, Jongseok Kim, Hyungseok Lee, Hyunsung Park 외

Change captioning is the task of identifying the change and describing it with a concise caption. Despite recent advancements, filtering out insignificant changes still remains as a challenge. Namely, images from dif…

MUNIChus: Multilingual News Image Captioning Benchmark

2026-03-11 · Yuji Chen, Alistair Plum, Hansi Hettiarachchi, Diptesh Kanojia 외 arxiv

The goal of news image captioning is to generate captions by integrating news article content with corresponding images, highlighting the relationship between textual context and visual elements. The majority of research…

Image Captioning

Multilingual Coreference Resolution via Cycle-Consistent Machine Translation

2026-06-03 · Adriana-Valentina Costache, Eduard Poesina, Silviu-Florin Gheorghe, Paul Irofti 외 arxiv

Coreference resolution is a core NLP task, having a broad range of downstream applications, e.g.~machine translation, question answering, document summarization, etc. While the task is well-studied in English, comparativ…

Document SummarizationCoreference ResolutionMachine TranslationQuestion Answering

AFRICAPTION: Establishing a New Paradigm for Image Captioning in African Languages

2025-10-20 · Mardiyyah Oduwole, Prince Mireku, Fatimo Adebanjo, Oluwatosin Olajide 외 arxiv

Multimodal AI research has overwhelmingly focused on high-resource languages, hindering the democratization of advancements in the field. To address this, we present AfriCaption, a comprehensive framework for multilingua…

Image Captioning