paper-with-me

Papers

Multilingual Image Corpus: Annotation Protocol

2021-09-01 · RANLP 2021 9 · Svetla Koeva

In this paper, we present work in progress aimed at the development of a new image dataset with annotated objects. The Multilingual Image Corpus consists of an ontology of visual objects (based on WordNet) and a collection of thematically related images annotated with segmentation masks and object classes. We identified 277 dominant classes and 1,037 parent and attribute classes, and grouped them into 10 thematic domains such as sport, medicine, education, food, security, etc. For the selected classes a large-scale web image search is being conducted in order to compile a substantial collection of high-quality copyright free images. The focus of the paper is the annotation protocol which we established to facilitate the annotation process: the Ontology of visual objects and the conventions for image selection and for object segmentation. The dataset is designed both for image classification and object detection and for semantic segmentation. In addition, the object annotations will be supplied with multilingual descriptions by using freely available wordnets.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Attributeimage-classificationImage ClassificationImage RetrievalObjectobject-detectionObject DetectionSegmentationSemantic Segmentation

Similar Papers 제목 키워드 기반

A Visually-Grounded Parallel Corpus with Phrase-to-Region Linking

2020-05-01 · LREC 2020 5 · Hideki Nakayama, Akihiro Tamura, Takashi Ninomiya

Visually-grounded natural language processing has become an important research direction in the past few years. However, majorities of the available cross-modal resources (e.g., image-caption datasets) are built in Engli…

Image CaptioningMachine TranslationMultimodal Machine TranslationTranslation

Multilingual Image Corpus – Towards a Multimodal and Multilingual Dataset

2022-06-01 · LREC 2022 6 · Svetla Koeva, Ivelina Stoyanova, Jordan Kralev

One of the processing tasks for large multimodal data streams is automatic image description (image classification, object segmentation and classification). Although the number and the diversity of image datasets is cons…

Caption Generationimage-classificationImage ClassificationImage Description+7

X-SRL: A Parallel Cross-Lingual Semantic Role Labeling Dataset

2020-10-05 · EMNLP 2020 11 · Angel Daza, Anette Frank

Even though SRL is researched for many languages, major improvements have mostly been obtained for English, for which more resources are available. In fact, existing multilingual SRL datasets contain disparate annotation…

Machine TranslationSemantic Role LabelingTranslation

Compiling Czech Parliamentary Stenographic Protocols into a Corpus

2020-05-01 · LREC 2020 5 · Barbora Hladka, Maty{\'a}{\v{s}} Kopp, Pavel Stra{\v{n}}{\'a}k

The Parliament of the Czech Republic consists of two chambers: the Chamber of Deputies (Lower House) and the Senate (Upper House). In our work, we focus on agenda and documents that relate to the Chamber of Deputies excl…

Lexical Coverage Evaluation of Large-scale Multilingual Semantic Lexicons for Twelve Languages

2016-05-01 · LREC 2016 5 · Scott Piao, Paul Rayson, Dawn Archer, Francesca Bianchi 외

The last two decades have seen the development of various semantic lexical resources such as WordNet (Miller, 1995) and the USAS semantic lexicon (Rayson et al., 2004), which have played an important role in the areas of…