paper-with-me

홈 › Papers

Uncovering divergent linguistic information in word embeddings with lessons for intrinsic and extrinsic evaluation

2018-09-06 · CONLL 2018 10 · Mikel Artetxe, Gorka Labaka, Iñigo Lopez-Gazpio, Eneko Agirre

Following the recent success of word embeddings, it has been argued that there is no such thing as an ideal representation for words, as different models tend to capture divergent and often mutually incompatible aspects like semantics/syntax and similarity/relatedness. In this paper, we show that each embedding model captures more information than directly apparent. A linear transformation that adjusts the similarity order of the model without any external resource can tailor it to achieve better results in those aspects, providing a new perspective on how embeddings encode divergent linguistic information. In addition, we explore the relation between intrinsic and extrinsic evaluation, as the effect of our transformations in downstream tasks is higher for unsupervised systems than for supervised ones.

📄 PDF Abstract BibTeX arXiv:1809.02094

Code (2)

artetxem/uncovec 공식 구현 pytorch
lgazpio/DAM_STS 공식 구현 pytorch

Tasks

Word Embeddings

Similar Papers 제목 키워드 기반

Field Embedding: A Unified Grain-Based Framework for Word Representation

2021-06-01 · NAACL 2021 4 · Junjie Luo, Xi Chen, Jichao Sun, Yuejia Xiang 외

Word representations empowered with additional linguistic information have been widely studied and proved to outperform traditional embeddings. Current methods mainly focus on learning embeddings for words while embeddin…

Word Embeddings

What company do words keep? Revisiting the distributional semantics of J.R. Firth & Zellig Harris

2022-05-16 · NAACL 2022 7 · Mikael Brunila, Jack LaViolette

The power of word embeddings is attributed to the linguistic theory that similar words will appear in similar contexts. This idea is specifically invoked by noting that "you shall know a word by the company it keeps," a …

Word Embeddings

What you can cram into a single vector: Probing sentence embeddings for linguistic properties

2018-05-03 · Alexis Conneau, German Kruszewski, Guillaume Lample, Loïc Barrault 외

Although much effort has recently been devoted to training high-quality sentence embeddings, we still have a poor understanding of what they are capturing. "Downstream" tasks, often based on sentence classification, are …

General ClassificationSentenceSentence ClassificationSentence Embeddings

What you can cram into a single \$\&!\#* vector: Probing sentence embeddings for linguistic properties

2018-07-01 · ACL 2018 7 · Alexis Conneau, German Kruszewski, Guillaume Lample, Lo{\"\i}c Barrault 외

Although much effort has recently been devoted to training high-quality sentence embeddings, we still have a poor understanding of what they are capturing. {``}Downstream{''} tasks, often based on sentence classification…

General ClassificationMachine TranslationSentenceSentence Classification+2

Uncovering Probabilistic Implications in Typological Knowledge Bases

2019-06-18 · ACL 2019 7 · Johannes Bjerva, Yova Kementchedjhieva, Ryan Cotterell, Isabelle Augenstein

The study of linguistic typology is rooted in the implications we find between linguistic features, such as the fact that languages with object-verb word ordering tend to have post-positions. Uncovering such implications…

Knowledge Base Population