paper-with-me

홈 › Papers

MultiSubs: A Large-scale Multimodal and Multilingual Dataset

2021-03-02 · LREC 2022 6 · Josiah Wang, Pranava Madhyastha, Josiel Figueiredo, Chiraag Lala, Lucia Specia

This paper introduces a large-scale multimodal and multilingual dataset that aims to facilitate research on grounding words to images in their contextual usage in language. The dataset consists of images selected to unambiguously illustrate concepts expressed in sentences from movie subtitles. The dataset is a valuable resource as (i) the images are aligned to text fragments rather than whole sentences; (ii) multiple images are possible for a text fragment and a sentence; (iii) the sentences are free-form and real-world like; (iv) the parallel texts are multilingual. We set up a fill-in-the-blank game for humans to evaluate the quality of the automatic image selection process of our dataset. We show the utility of the dataset on two automatic tasks: (i) fill-in-the-blank; (ii) lexical translation. Results of the human evaluation and automatic models demonstrate that images can be a useful complement to the textual context. The dataset will benefit research on visual grounding of words especially in the context of free-form sentences, and can be obtained from https://doi.org/10.5281/zenodo.5034604 under a Creative Commons licence.

📄 PDF Abstract BibTeX arXiv:2103.01910

Code (1)

josiahwang/multisubs-eval 공식 구현

Tasks

Multimodal Lexical TranslationMultimodal Text PredictionSentenceTranslation

Similar Papers 제목 키워드 기반

mOSCAR: A Large-scale Multilingual and Multimodal Document-level Corpus

2024-06-13 · Matthieu Futeral, Armel Zebaze, Pedro Ortiz Suarez, Julien Abadji 외

Multimodal Large Language Models (mLLMs) are trained on a large amount of text-image data. While most mLLMs are trained on caption-like data only, Alayrac et al. [2022] showed that additionally training them on interleav…

Few-Shot LearningIn-Context Learning

Grounding Multilingual Multimodal LLMs With Cultural Knowledge

2025-08-10 · Jean de Dieu Nyandwi, Yueqi Song, Simran Khanuja, Graham Neubig arxiv

Multimodal Large Language Models excel in high-resource settings, but often misinterpret long-tail cultural entities and underperform in low-resource languages. To address this gap, we propose a data-centric approach tha…

Visual Question Answering

2M-NER: Contrastive Learning for Multilingual and Multimodal NER with Language and Modal Fusion

2024-04-26 · Dongsheng Wang, Xiaoqin Feng, Zeming Liu, Chuan Wang

Named entity recognition (NER) is a fundamental task in natural language processing that involves identifying and classifying entities in sentences into pre-defined types. It plays a crucial role in various research fiel…

Contrastive LearningEntity Linkingnamed-entity-recognitionNamed Entity Recognition+6

CMU-MOSEAS: A Multimodal Language Dataset for Spanish, Portuguese, German and French

2020-11-01 · EMNLP 2020 11 · AmirAli Bagher Zadeh, Yansheng Cao, Simon Hessner, Paul Pu Liang 외

Modeling multimodal language is a core research area in natural language processing. While languages such as English have relatively large multimodal language resources, other widely spoken languages across the globe hav…

Kaleidoscope: In-language Exams for Massively Multilingual Vision Evaluation

2025-04-09 · Israfel Salazar, Manuel Fernández Burda, Shayekh Bin Islam, Arshia Soltani Moakhar 외

The evaluation of vision-language models (VLMs) has mainly relied on English-language benchmarks, leaving significant gaps in both multilingual and multicultural coverage. While multilingual benchmarks have expanded, bot…

Multiple-choice