COCO-CN for Cross-Lingual Image Tagging, Captioning and Retrieval
This paper contributes to cross-lingual image annotation and retrieval in terms of data and baseline methods. We propose COCO-CN, a novel dataset enriching MS-COCO with manually written Chinese sentences and tags. For more effective annotation acquisition, we develop a recommendation-assisted collective annotation system, automatically providing an annotator with several tags and sentences deemed to be relevant with respect to the pictorial content. Having 20,342 images annotated with 27,218 Chinese sentences and 70,993 tags, COCO-CN is currently the largest Chinese-English dataset that provides a unified and challenging platform for cross-lingual image tagging, captioning and retrieval. We develop conceptually simple yet effective methods per task for learning from cross-lingual resources. Extensive experiments on the three tasks justify the viability of the proposed dataset and methods. Data and code are publicly available at https://github.com/li-xirong/coco-cn
Code (2)
Tasks
RetrievalSimilar Papers 제목 키워드 기반
Towards Zero-shot Cross-lingual Image Retrieval and Tagging
There has been a recent spike in interest in multi-modal Language and Vision problems. On the language side, most of these models primarily focus on English since most multi-modal datasets are monolingual. We try to brid…
Image RetrievalRetrievalEmbedded Heterogeneous Attention Transformer for Cross-lingual Image Captioning
Cross-lingual image captioning is a challenging task that requires addressing both cross-lingual and cross-modal obstacles in multimedia analysis. The crucial issue in this task is to model the global and the local match…
Image CaptioningFlorenz: Scaling Laws for Systematic Generalization in Vision-Language Models
Cross-lingual transfer enables vision-language models (VLMs) to perform vision tasks in various languages with training data only in one language. Current approaches rely on large pre-trained multilingual language models…
Cross-Lingual TransferImage CaptioningLarge Language ModelMachine Translation+3CAPTION: Correction by Analyses, POS-Tagging and Interpretation of Objects using only Nouns
Recently, Deep Learning (DL) methods have shown an excellent performance in image captioning and visual question answering. However, despite their performance, DL methods do not learn the semantics of the words that are …
Image Captioningobject-detectionObject DetectionPOS+4Multi-view and Cross-view Brain Decoding
Can we build multi-view decoders that can decode concepts from brain recordings corresponding to any view (picture, sentence, word cloud) of stimuli? Can we build a system that can use brain recordings to automatically d…
Brain DecodingImage CaptioningKeyword ExtractionSentence