paper-with-me

Papers

MultiNews: A Web collection of an Aligned Multimodal and Multilingual Corpus

2017-11-01 · WS 2017 11 · Haithem Afli, Pintu Lohar, Andy Way

Integrating Natural Language Processing (NLP) and computer vision is a promising effort. However, the applicability of these methods directly depends on the availability of a specific multimodal data that includes images and texts. In this paper, we present a collection of a Multimodal corpus of comparable texts and their images in 9 languages from the web news articles of Euronews website. This corpus has found widespread use in the NLP community in Multilingual and multimodal tasks. Here, we focus on its acquisition of the images and text data and their multilingual alignment.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

ArticlesContent-Based Image RetrievalImage RetrievalMachine TranslationMultimodal Machine Translation

Similar Papers 제목 키워드 기반

Multilingual Image Corpus – Towards a Multimodal and Multilingual Dataset

2022-06-01 · LREC 2022 6 · Svetla Koeva, Ivelina Stoyanova, Jordan Kralev

One of the processing tasks for large multimodal data streams is automatic image description (image classification, object segmentation and classification). Although the number and the diversity of image datasets is cons…

Caption Generationimage-classificationImage ClassificationImage Description+7

The AMARA Corpus: Building Parallel Language Resources for the Educational Domain

2014-05-01 · LREC 2014 5 · Ahmed Abdelali, Francisco Guzman, Hassan Sajjad, Stephan Vogel

This paper presents the AMARA corpus of on-line educational content: a new parallel corpus of educational video subtitles, multilingually aligned for 20 languages, i.e. 20 monolingual corpora and 190 parallel corpora. Th…

Machine TranslationTranslation

Automatic Translation Alignment for Ancient Greek and Latin

2022-06-01 · LT4HALA (LREC) 2022 6 · Tariq Yousef, Chiara Palladino, David J. Wright, Monica Berti

This paper presents the results of automatic translation alignment experiments on a corpus of texts in Ancient Greek translated into Latin. We used a state-of-the-art alignment workflow based on a contextualized multilin…

Language ModelingLanguage ModellingTranslation

Phoenix-VL 1.5 Medium Technical Report

2026-05-11 · Team Phoenix, :, Arka Ray, Askar Ali Mohamed Jawad 외 arxiv

We introduce Phoenix-VL 1.5 Medium, a 123B-parameter natively multimodal and multilingual foundation model, adapted to regional languages and the Singapore context. Developed as a sovereign AI asset, it demonstrates that…

Domain Adaptation

From Unaligned to Aligned: Scaling Multilingual LLMs with Multi-Way Parallel Corpora

2025-05-20 · Yingli Shen, Wen Lai, Shuo Wang, Kangyang Luo 외

Continued pretraining and instruction tuning on large-scale multilingual data have proven to be effective in scaling large language models (LLMs) to low-resource languages. However, the unaligned nature of such data limi…