paper-with-me

홈 › Papers

Crosslingual and Multilingual Construction of Syntax-Based Vector Space Models

2014-01-01 · TACL 2014 1 · Jason Utt, Sebastian Pad{\'o}

Syntax-based distributional models of lexical semantics provide a flexible and linguistically adequate representation of co-occurrence information. However, their construction requires large, accurately parsed corpora, which are unavailable for most languages. In this paper, we develop a number of methods to overcome this obstacle. We describe (a) a crosslingual approach that constructs a syntax-based model for a new language requiring only an English resource and a translation lexicon; and (b) multilingual approaches that combine crosslingual with monolingual information, subject to availability. We evaluate on two lexical semantic benchmarks in German and Croatian. We find that the models exhibit complementary profiles: crosslingual models yield higher accuracies while monolingual models provide better coverage. In addition, we show that simple multilingual models can successfully combine their strengths.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Translation

Similar Papers 제목 키워드 기반

Multilingual Training of Crosslingual Word Embeddings

2017-04-01 · EACL 2017 4 · Long Duong, Hiroshi Kanayama, Tengfei Ma, Steven Bird 외

Crosslingual word embeddings represent lexical items from different languages using the same vector space, enabling crosslingual transfer. Most prior work constructs embeddings for a pair of languages, with English on on…

Bilingual Lexicon InductionDependency ParsingDocument ClassificationGeneral Classification+6

Embedding Learning Through Multilingual Concept Induction

2018-01-21 · ACL 2018 7 · Philipp Dufter, Mengjie Zhao, Martin Schmitt, Alexander Fraser 외

We present a new method for estimating vector space representations of words: embedding learning by concept induction. We test this method on a highly parallel corpus and learn semantic representations of words in 1259 d…

Sentiment AnalysisWord Similarity

Crosslingual Document Embedding as Reduced-Rank Ridge Regression

2019-04-08 · Martin Josifoski, Ivan S. Paskov, Hristo S. Paskov, Martin Jaggi 외

There has recently been much interest in extending vector-based word representations to multiple languages, such that words can be compared across languages. In this paper, we shift the focus from words to documents and …

Document EmbeddingregressionRetrievalSentence

Crosslingual Capabilities and Knowledge Barriers in Multilingual Large Language Models

2024-06-23 · Lynn Chua, Badih Ghazi, Yangsibo Huang, Pritish Kamath 외

Large language models (LLMs) are typically multilingual due to pretraining on diverse multilingual corpora. But can these models relate corresponding concepts across languages, effectively being crosslingual? This study …

Machine TranslationMMLUTransfer Learning

Multilingual and crosslingual speech recognition using phonological-vector based phone embeddings

2021-07-11 · Chengrui Zhu, Keyu An, Huahuan Zheng, Zhijian Ou

The use of phonological features (PFs) potentially allows language-specific phones to remain linked in training, which is highly desirable for information sharing for multilingual and crosslingual speech recognition meth…

speech-recognitionSpeech Recognition