paper-with-me

홈 › Papers

COCO-CN for Cross-Lingual Image Tagging, Captioning and Retrieval

2018-05-22 · Xirong Li, Chaoxi Xu, Xiaoxu Wang, Weiyu Lan, Zhengxiong Jia, Gang Yang, Jieping Xu

This paper contributes to cross-lingual image annotation and retrieval in terms of data and baseline methods. We propose COCO-CN, a novel dataset enriching MS-COCO with manually written Chinese sentences and tags. For more effective annotation acquisition, we develop a recommendation-assisted collective annotation system, automatically providing an annotator with several tags and sentences deemed to be relevant with respect to the pictorial content. Having 20,342 images annotated with 27,218 Chinese sentences and 70,993 tags, COCO-CN is currently the largest Chinese-English dataset that provides a unified and challenging platform for cross-lingual image tagging, captioning and retrieval. We develop conceptually simple yet effective methods per task for learning from cross-lingual resources. Extensive experiments on the three tasks justify the viability of the proposed dataset and methods. Data and code are publicly available at https://github.com/li-xirong/coco-cn

📄 PDF Abstract BibTeX arXiv:1805.08661

Code (2)

li-xirong/coco-cn 공식 구현
evanmiltenburg/COCO-CN-Results-Viewer

Tasks

Retrieval

Similar Papers 제목 키워드 기반

Towards Zero-shot Cross-lingual Image Retrieval and Tagging

2021-09-15 · Pranav Aggarwal, Ritiz Tambi, Ajinkya Kale

There has been a recent spike in interest in multi-modal Language and Vision problems. On the language side, most of these models primarily focus on English since most multi-modal datasets are monolingual. We try to brid…

Image RetrievalRetrieval

Embedded Heterogeneous Attention Transformer for Cross-lingual Image Captioning

2023-07-19 · Zijie Song, Zhenzhen Hu, Yuanen Zhou, Ye Zhao 외

Cross-lingual image captioning is a challenging task that requires addressing both cross-lingual and cross-modal obstacles in multimedia analysis. The crucial issue in this task is to model the global and the local match…

Image Captioning

Florenz: Scaling Laws for Systematic Generalization in Vision-Language Models

2025-03-12 · Julian Spravil, Sebastian Houben, Sven Behnke

Cross-lingual transfer enables vision-language models (VLMs) to perform vision tasks in various languages with training data only in one language. Current approaches rely on large pre-trained multilingual language models…

Cross-Lingual TransferImage CaptioningLarge Language ModelMachine Translation+3

CAPTION: Correction by Analyses, POS-Tagging and Interpretation of Objects using only Nouns

2020-10-02 · Leonardo Anjoletto Ferreira, Douglas De Rizzo Meneghetti, Paulo Eduardo Santos

Recently, Deep Learning (DL) methods have shown an excellent performance in image captioning and visual question answering. However, despite their performance, DL methods do not learn the semantics of the words that are …

Image Captioningobject-detectionObject DetectionPOS+4

Multi-view and Cross-view Brain Decoding

2022-10-01 · COLING 2022 10 · Subba Reddy Oota, Jashn Arora, Manish Gupta, Raju S. Bapi

Can we build multi-view decoders that can decode concepts from brain recordings corresponding to any view (picture, sentence, word cloud) of stimuli? Can we build a system that can use brain recordings to automatically d…

Brain DecodingImage CaptioningKeyword ExtractionSentence