Multilingual Image Corpus: Annotation Protocol
In this paper, we present work in progress aimed at the development of a new image dataset with annotated objects. The Multilingual Image Corpus consists of an ontology of visual objects (based on WordNet) and a collection of thematically related images annotated with segmentation masks and object classes. We identified 277 dominant classes and 1,037 parent and attribute classes, and grouped them into 10 thematic domains such as sport, medicine, education, food, security, etc. For the selected classes a large-scale web image search is being conducted in order to compile a substantial collection of high-quality copyright free images. The focus of the paper is the annotation protocol which we established to facilitate the annotation process: the Ontology of visual objects and the conventions for image selection and for object segmentation. The dataset is designed both for image classification and object detection and for semantic segmentation. In addition, the object annotations will be supplied with multilingual descriptions by using freely available wordnets.
Code (0)
등록된 구현이 없습니다.
Tasks
Attributeimage-classificationImage ClassificationImage RetrievalObjectobject-detectionObject DetectionSegmentationSemantic SegmentationSimilar Papers 제목 키워드 기반
A Visually-Grounded Parallel Corpus with Phrase-to-Region Linking
Visually-grounded natural language processing has become an important research direction in the past few years. However, majorities of the available cross-modal resources (e.g., image-caption datasets) are built in Engli…
Image CaptioningMachine TranslationMultimodal Machine TranslationTranslationMultilingual Image Corpus – Towards a Multimodal and Multilingual Dataset
One of the processing tasks for large multimodal data streams is automatic image description (image classification, object segmentation and classification). Although the number and the diversity of image datasets is cons…
Caption Generationimage-classificationImage ClassificationImage Description+7X-SRL: A Parallel Cross-Lingual Semantic Role Labeling Dataset
Even though SRL is researched for many languages, major improvements have mostly been obtained for English, for which more resources are available. In fact, existing multilingual SRL datasets contain disparate annotation…
Machine TranslationSemantic Role LabelingTranslationCompiling Czech Parliamentary Stenographic Protocols into a Corpus
The Parliament of the Czech Republic consists of two chambers: the Chamber of Deputies (Lower House) and the Senate (Upper House). In our work, we focus on agenda and documents that relate to the Chamber of Deputies excl…
Lexical Coverage Evaluation of Large-scale Multilingual Semantic Lexicons for Twelve Languages
The last two decades have seen the development of various semantic lexical resources such as WordNet (Miller, 1995) and the USAS semantic lexicon (Rayson et al., 2004), which have played an important role in the areas of…