paper-with-me

홈 › Papers

Huge Automatically Extracted Training Sets for Multilingual Word Sense Disambiguation

2018-05-12 · Tommaso Pasini, Francesco Maria Elia, Roberto Navigli

We release to the community six large-scale sense-annotated datasets in multiple language to pave the way for supervised multilingual Word Sense Disambiguation. Our datasets cover all the nouns in the English WordNet and their translations in other languages for a total of millions of sense-tagged sentences. Experiments prove that these corpora can be effectively used as training sets for supervised WSD systems, surpassing the state of the art for low-resourced languages and providing competitive results for English, where manually annotated training sets are accessible. The data is available at trainomatic.org.

📄 PDF Abstract BibTeX arXiv:1805.04685

Code (0)

등록된 구현이 없습니다.

Tasks

Word Sense Disambiguation

Similar Papers 제목 키워드 기반

Huge Automatically Extracted Training-Sets for Multilingual Word SenseDisambiguation

2018-05-01 · LREC 2018 5 · Tommaso Pasini, Francesco Elia, Roberto Navigli
Question AnsweringSemantic ParsingWord Sense Disambiguation

Code-switched inspired losses for spoken dialog representations

2021-11-01 · EMNLP 2021 11 · Pierre Colombo, Emile Chapuis, Matthieu Labeau, Chloé Clavel

Spoken dialogue systems need to be able to handle both multiple languages and multilinguality inside a conversation (e.g in case of code-switching). In this work, we introduce new pretraining losses tailored to learn gen…

RetrievalSpoken Dialogue Systems

Code-switched inspired losses for generic spoken dialog representations

2021-08-27 · Emile Chapuis, Pierre Colombo, Matthieu Labeau, Chloe Clavel

Spoken dialog systems need to be able to handle both multiple languages and multilinguality inside a conversation (\textit{e.g} in case of code-switching). In this work, we introduce new pretraining losses tailored to le…

Retrieval

Vastextures: Vast repository of textures and PBR materials extracted from real-world images using unsupervised methods

2024-06-24 · Sagi Eppel

Vastextures is a vast repository of 500,000 textures and PBR materials extracted from real-world images using an unsupervised process. The extracted materials and textures are extremely diverse and cover a vast range of …

Bilingual Language Modeling, A transfer learning technique for Roman Urdu

2021-02-22 · Usama Khalid, Mirza Omer Beg, Muhammad Umair Arshad

Pretrained language models are now of widespread use in Natural Language Processing. Despite their success, applying them to Low Resource languages is still a huge challenge. Although Multilingual models hold great promi…

Cross-Lingual TransferLanguage ModelingLanguage ModellingMasked Language Modeling+1