paper-with-me

홈 › Papers

A Simple Regularization-based Algorithm for Learning Cross-Domain Word Embeddings

2019-02-01 · EMNLP 2017 9 · Wei Yang, Wei Lu, Vincent W. Zheng

Learning word embeddings has received a significant amount of attention recently. Often, word embeddings are learned in an unsupervised manner from a large collection of text. The genre of the text typically plays an important role in the effectiveness of the resulting embeddings. How to effectively train word embedding models using data from different domains remains a problem that is underexplored. In this paper, we present a simple yet effective method for learning word embeddings based on text from different domains. We demonstrate the effectiveness of our approach through extensive experiments on various down-stream NLP tasks.

📄 PDF Abstract BibTeX arXiv:1902.00184

Code (0)

등록된 구현이 없습니다.

Tasks

Learning Word EmbeddingsWord Embeddings

Similar Papers 제목 키워드 기반

Subword Regularization: Improving Neural Network Translation Models with Multiple Subword Candidates

2018-04-29 · ACL 2018 7 · Taku Kudo

Subword units are an effective way to alleviate the open vocabulary problems in neural machine translation (NMT). While sentences are usually converted into unique subword sequences, subword segmentation is potentially a…

Language ModelingLanguage ModellingMachine TranslationNMT+2

MASKER: Masked Keyword Regularization for Reliable Text Classification

2020-12-17 · Seung Jun Moon, Sangwoo Mo, Kimin Lee, Jaeho Lee 외

Pre-trained language models have achieved state-of-the-art accuracies on various text classification tasks, e.g., sentiment analysis, natural language inference, and semantic textual similarity. However, the reliability …

ClassificationDomain GeneralizationGeneral ClassificationNatural Language Inference+5

BPE-Dropout: Simple and Effective Subword Regularization

2019-10-29 · ACL 2020 6 · Ivan Provilkov, Dmitrii Emelianenko, Elena Voita

Subword segmentation is widely used to address the open vocabulary problem in machine translation. The dominant approach to subword segmentation is Byte Pair Encoding (BPE), which keeps the most frequent words intact whi…

Machine TranslationSegmentationTranslation

Neural Chinese Word Segmentation with Lexicon and Unlabeled Data via Posterior Regularization

2019-04-26 · Junxin Liu, Fangzhao Wu, Chuhan Wu, Yongfeng Huang 외

Existing methods for CWS usually rely on a large number of labeled sentences to train word segmentation models, which are expensive and time-consuming to annotate. Luckily, the unlabeled data is usually easy to collect a…

Chinese Word SegmentationSegmentation

Beyond the Deep Metric Learning: Enhance the Cross-Modal Matching with Adversarial Discriminative Domain Regularization

2020-10-23 · Li Ren, Kai Li, Liqiang Wang, Kien Hua

Matching information across image and text modalities is a fundamental challenge for many applications that involve both vision and natural language processing. The objective is to find efficient similarity metrics to co…

Metric LearningSentence