Towards Annotation-Free Evaluation of Cross-Lingual Image Captioning
Cross-lingual image captioning, with its ability to caption an unlabeled image in a target language other than English, is an emerging topic in the multimedia field. In order to save the precious human resource from re-writing reference sentences per target language, in this paper we make a brave attempt towards annotation-free evaluation of cross-lingual image captioning. Depending on whether we assume the availability of English references, two scenarios are investigated. For the first scenario with the references available, we propose two metrics, i.e., WMDRel and CLinRel. WMDRel measures the semantic relevance between a model-generated caption and machine translation of an English reference using their Word Mover's Distance. By projecting both captions into a deep visual feature space, CLinRel is a visual-oriented cross-lingual relevance measure. As for the second scenario, which has zero reference and is thus more challenging, we propose CMedRel to compute a cross-media relevance between the generated caption and the image content, in the same visual feature space as used by CLinRel. The promising results show high potential of the new metrics for evaluation with no need of references in the target language.
Code (0)
등록된 구현이 없습니다.
Tasks
Image CaptioningMachine TranslationTranslationSimilar Papers 제목 키워드 기반
ANHALTEN: Cross-Lingual Transfer for German Token-Level Reference-Free Hallucination Detection
Research on token-level reference-free hallucination detection has predominantly focused on English, primarily due to the scarcity of robust datasets in other languages. This has hindered systematic investigations into t…
Cross-Lingual TransferHallucinationMultilingual Image Corpus: Annotation Protocol
In this paper, we present work in progress aimed at the development of a new image dataset with annotated objects. The Multilingual Image Corpus consists of an ontology of visual objects (based on WordNet) and a collecti…
Attributeimage-classificationImage ClassificationImage Retrieval+5Multilingual Projection for Parsing Truly Low-Resource Languages
We propose a novel approach to cross-lingual part-of-speech tagging and dependency parsing for truly low-resource languages. Our annotation projection-based approach yields tagging and parsing models for over 100 languag…
Cross-Lingual TransferDependency ParsingPart-Of-Speech TaggingTransfer LearningRECSA: Resource for Evaluating Cross-lingual Semantic Annotation
In recent years large repositories of structured knowledge (DBpedia, Freebase, YAGO) have become a valuable resource for language technologies, especially for the automatic aggregation of knowledge from textual data. One…
ArticlesMachine TranslationSpatialVOC2K: A Multilingual Dataset of Images with Annotations and Features for Spatial Relations between Objects
We present SpatialVOC2K, the first multilingual image dataset with spatial relation annotations and object features for image-to-text generation, built using 2,026 images from the PASCAL VOC2008 dataset. The dataset inco…
Image to textObjectText Generation