paper-with-me

Papers

Learning Multilingual Word Representations using a Bag-of-Words Autoencoder

2014-01-08 · Stanislas Lauly, Alex Boulanger, Hugo Larochelle

Recent work on learning multilingual word representations usually relies on the use of word-level alignements (e.g. infered with the help of GIZA++) between translated sentences, in order to align the word embeddings in different languages. In this workshop paper, we investigate an autoencoder model for learning multilingual word representations that does without such word-level alignements. The autoencoder is trained to reconstruct the bag-of-word representation of given sentence from an encoded representation extracted from its translation. We evaluate our approach on a multilingual document classification task, where labeled data is available only for one language (e.g. English) while classification must be performed in a different language (e.g. French). In our experiments, we observe that our method compares favorably with a previously proposed method that exploits word-level alignments to learn word representations.

📄 PDF Abstract BibTeX arXiv:1401.1803

Code (0)

등록된 구현이 없습니다.

Tasks

Document ClassificationGeneral ClassificationSentenceTranslationWord Embeddings

Methods 이 논문이 사용한 방법론

Solana Customer Service Number +1-833-534-1729 설명 없음

Similar Papers 제목 키워드 기반

Multilingual Vector Representations of Words, Sentences, and Documents

2017-11-01 · IJCNLP 2017 11 · Gerard de Melo

Neural vector representations are now ubiquitous in all subfields of natural language processing and text mining. While methods such as word2vec and GloVe are well-known, this tutorial focuses on multilingual and cross-l…

Knowledge Graphs

Embedding Learning Through Multilingual Concept Induction

2018-01-21 · ACL 2018 7 · Philipp Dufter, Mengjie Zhao, Martin Schmitt, Alexander Fraser 외

We present a new method for estimating vector space representations of words: embedding learning by concept induction. We test this method on a highly parallel corpus and learn semantic representations of words in 1259 d…

Sentiment AnalysisWord Similarity

Gender Bias in Multilingual Embeddings and Cross-Lingual Transfer

2020-05-02 · ACL 2020 6 · Jieyu Zhao, Subhabrata Mukherjee, Saghar Hosseini, Kai-Wei Chang 외

Multilingual representations embed words from many languages into a single semantic space such that words with similar meanings are close to each other regardless of the language. These embeddings have been widely used i…

Cross-Lingual TransferTransfer Learning

Improved acoustic word embeddings for zero-resource languages using multilingual transfer

2020-06-02 · Herman Kamper, Yevgen Matusevych, Sharon Goldwater

Acoustic word embeddings are fixed-dimensional representations of variable-length speech segments. Such embeddings can form the basis for speech search, indexing and discovery systems when conventional speech recognition…

speech-recognitionSpeech RecognitionWord Embeddings

Multilingual Factor Analysis

2019-05-14 · ACL 2019 7 · Francisco Vargas, Kamen Brestnichki, Alex Papadopoulos-Korfiatis, Nils Hammerla

In this work we approach the task of learning multilingual word representations in an offline manner by fitting a generative latent variable model to a multilingual dictionary. We model equivalent words in different lang…