Combining Discourse Markers and Cross-lingual Embeddings for Synonym--Antonym Classification
It is well-known that distributional semantic approaches have difficulty in distinguishing between synonyms and antonyms (Grefenstette, 1992; Pad{\'o} and Lapata, 2003). Recent work has shown that supervision available in English for this task (e.g., lexical resources) can be transferred to other languages via cross-lingual word embeddings. However, this kind of transfer misses monolingual distributional information available in a target language, such as contrast relations that are indicative of antonymy (e.g. hot ... while ... cold). In this work, we improve the transfer by exploiting monolingual information, expressed in the form of co-occurrences with discourse markers that convey contrast. Our approach makes use of less than a dozen markers, which can easily be obtained for many languages. Compared to a baseline using only cross-lingual embeddings, we show absolute improvements of 4{--}10{\%} F1-score in Vietnamese and Hindi.
Code (0)
등록된 구현이 없습니다.
Tasks
Cross-Lingual Word EmbeddingsGeneral ClassificationWord EmbeddingsSimilar Papers 제목 키워드 기반
Towards Using Machine Translation Techniques to Induce Multilingual Lexica of Discourse Markers
Discourse markers are universal linguistic events subject to language variation. Although an extensive literature has already reported language specific traits of these events, little has been said on their cross-languag…
Machine TranslationSentenceTranslationTracing variation in discourse connectives in translation and interpreting through neural semantic spaces
In the present paper, we explore lexical contexts of discourse markers in translation and interpreting on the basis of word embeddings. Our special interest is on contextual variation of the same discourse markers in (wr…
TranslationWord EmbeddingsISO-based Annotated Multilingual Parallel Corpus for Discourse Markers
Discourse markers carry information about the discourse structure and organization, and also signal local dependencies or epistemological stance of speaker. They provide instructions on how to interpret the discourse, an…
Mining Discourse Markers for Unsupervised Sentence Representation Learning
Current state of the art systems in NLP heavily rely on manually annotated datasets, which are expensive to construct. Very little work adequately exploits unannotated data -- such as discourse markers between sentences …
Relation ClassificationRepresentation LearningSentenceSentence EmbeddingsInducing Discourse Marker Inventories from Lexical Knowledge Graphs
Discourse marker inventories are important tools for the development of both discourse parsers and corpora with discourse annotations. In this paper we explore the potential of massively multilingual lexical knowledge gr…
Knowledge Graphs