paper-with-me

Papers

Multilingual Image Corpus – Towards a Multimodal and Multilingual Dataset

2022-06-01 · LREC 2022 6 · Svetla Koeva, Ivelina Stoyanova, Jordan Kralev

One of the processing tasks for large multimodal data streams is automatic image description (image classification, object segmentation and classification). Although the number and the diversity of image datasets is constantly expanding, still there is a huge demand for more datasets in terms of variety of domains and object classes covered. The goal of the project Multilingual Image Corpus (MIC 21) is to provide a large image dataset with annotated objects and object descriptions in 24 languages. The Multilingual Image Corpus consists of an Ontology of visual objects (based on WordNet) and a collection of thematically related images whose objects are annotated with segmentation masks and labels describing the ontology classes. The dataset is designed both for image classification and object detection and for semantic segmentation. The main contributions of our work are: a) the provision of large collection of high quality copyright-free images; b) the formulation of the Ontology of visual objects based on WordNet noun hierarchies; c) the precise manual correction of automatic object segmentation within the images and the annotation of object classes; and d) the association of objects and images with extended multilingual descriptions based on WordNet inner- and interlingual relations. The dataset can be used also for multilingual image caption generation, image-to-text alignment and automatic question answering for images and videos.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Caption Generationimage-classificationImage ClassificationImage DescriptionImage to textObjectobject-detectionObject DetectionQuestion AnsweringSegmentationSemantic Segmentation

Similar Papers 제목 키워드 기반

MultiNews: A Web collection of an Aligned Multimodal and Multilingual Corpus

2017-11-01 · WS 2017 11 · Haithem Afli, Pintu Lohar, Andy Way

Integrating Natural Language Processing (NLP) and computer vision is a promising effort. However, the applicability of these methods directly depends on the availability of a specific multimodal data that includes images…

ArticlesContent-Based Image RetrievalImage RetrievalMachine Translation+1

A Visually-Grounded Parallel Corpus with Phrase-to-Region Linking

2020-05-01 · LREC 2020 5 · Hideki Nakayama, Akihiro Tamura, Takashi Ninomiya

Visually-grounded natural language processing has become an important research direction in the past few years. However, majorities of the available cross-modal resources (e.g., image-caption datasets) are built in Engli…

Image CaptioningMachine TranslationMultimodal Machine TranslationTranslation

mOSCAR: A Large-scale Multilingual and Multimodal Document-level Corpus

2024-06-13 · Matthieu Futeral, Armel Zebaze, Pedro Ortiz Suarez, Julien Abadji 외

Multimodal Large Language Models (mLLMs) are trained on a large amount of text-image data. While most mLLMs are trained on caption-like data only, Alayrac et al. [2022] showed that additionally training them on interleav…

Few-Shot LearningIn-Context Learning

Bridge Correlational Neural Networks for Multilingual Multimodal Representation Learning

2015-10-13 · NAACL 2016 6 · Janarthanan Rajendran, Mitesh M. Khapra, Sarath Chandar, Balaraman Ravindran

Recently there has been a lot of interest in learning common representations for multiple views of data. Typically, such common representations are learned using a parallel corpus between the two views (say, 1M images an…

Document ClassificationRepresentation LearningRetrievalTransfer Learning

Multilingual-To-Multimodal (M2M): Unlocking New Languages with Monolingual Text

2026-01-15 · Piyush Singh Pasi arxiv

Multimodal models excel in English, supported by abundant image-text and audio-text data, but performance drops sharply for other languages due to limited multilingual multimodal resources. Existing solutions rely on mac…

Text-to-Image GenerationMachine TranslationImage RetrievalText Retrieval