Huge Automatically Extracted Training Sets for Multilingual Word Sense Disambiguation
We release to the community six large-scale sense-annotated datasets in multiple language to pave the way for supervised multilingual Word Sense Disambiguation. Our datasets cover all the nouns in the English WordNet and their translations in other languages for a total of millions of sense-tagged sentences. Experiments prove that these corpora can be effectively used as training sets for supervised WSD systems, surpassing the state of the art for low-resourced languages and providing competitive results for English, where manually annotated training sets are accessible. The data is available at trainomatic.org.
Code (0)
등록된 구현이 없습니다.
Tasks
Word Sense DisambiguationSimilar Papers 제목 키워드 기반
Huge Automatically Extracted Training-Sets for Multilingual Word SenseDisambiguation
Code-switched inspired losses for spoken dialog representations
Spoken dialogue systems need to be able to handle both multiple languages and multilinguality inside a conversation (e.g in case of code-switching). In this work, we introduce new pretraining losses tailored to learn gen…
RetrievalSpoken Dialogue SystemsCode-switched inspired losses for generic spoken dialog representations
Spoken dialog systems need to be able to handle both multiple languages and multilinguality inside a conversation (\textit{e.g} in case of code-switching). In this work, we introduce new pretraining losses tailored to le…
RetrievalVastextures: Vast repository of textures and PBR materials extracted from real-world images using unsupervised methods
Vastextures is a vast repository of 500,000 textures and PBR materials extracted from real-world images using an unsupervised process. The extracted materials and textures are extremely diverse and cover a vast range of …
Bilingual Language Modeling, A transfer learning technique for Roman Urdu
Pretrained language models are now of widespread use in Natural Language Processing. Despite their success, applying them to Low Resource languages is still a huge challenge. Although Multilingual models hold great promi…
Cross-Lingual TransferLanguage ModelingLanguage ModellingMasked Language Modeling+1