MultiNews: A Web collection of an Aligned Multimodal and Multilingual Corpus
Integrating Natural Language Processing (NLP) and computer vision is a promising effort. However, the applicability of these methods directly depends on the availability of a specific multimodal data that includes images and texts. In this paper, we present a collection of a Multimodal corpus of comparable texts and their images in 9 languages from the web news articles of Euronews website. This corpus has found widespread use in the NLP community in Multilingual and multimodal tasks. Here, we focus on its acquisition of the images and text data and their multilingual alignment.
Code (0)
등록된 구현이 없습니다.
Tasks
ArticlesContent-Based Image RetrievalImage RetrievalMachine TranslationMultimodal Machine TranslationSimilar Papers 제목 키워드 기반
Multilingual Image Corpus – Towards a Multimodal and Multilingual Dataset
One of the processing tasks for large multimodal data streams is automatic image description (image classification, object segmentation and classification). Although the number and the diversity of image datasets is cons…
Caption Generationimage-classificationImage ClassificationImage Description+7The AMARA Corpus: Building Parallel Language Resources for the Educational Domain
This paper presents the AMARA corpus of on-line educational content: a new parallel corpus of educational video subtitles, multilingually aligned for 20 languages, i.e. 20 monolingual corpora and 190 parallel corpora. Th…
Machine TranslationTranslationAutomatic Translation Alignment for Ancient Greek and Latin
This paper presents the results of automatic translation alignment experiments on a corpus of texts in Ancient Greek translated into Latin. We used a state-of-the-art alignment workflow based on a contextualized multilin…
Language ModelingLanguage ModellingTranslationPhoenix-VL 1.5 Medium Technical Report
We introduce Phoenix-VL 1.5 Medium, a 123B-parameter natively multimodal and multilingual foundation model, adapted to regional languages and the Singapore context. Developed as a sovereign AI asset, it demonstrates that…
Domain AdaptationFrom Unaligned to Aligned: Scaling Multilingual LLMs with Multi-Way Parallel Corpora
Continued pretraining and instruction tuning on large-scale multilingual data have proven to be effective in scaling large language models (LLMs) to low-resource languages. However, the unaligned nature of such data limi…