paper-with-me

Papers

Combining Pretrained High-Resource Embeddings and Subword Representations for Low-Resource Languages

2020-03-09 · Machel Reid, Edison Marrese-Taylor, Yutaka Matsuo

The contrast between the need for large amounts of data for current Natural Language Processing (NLP) techniques, and the lack thereof, is accentuated in the case of African languages, most of which are considered low-resource. To help circumvent this issue, we explore techniques exploiting the qualities of morphologically rich languages (MRLs), while leveraging pretrained word vectors in well-resourced languages. In our exploration, we show that a meta-embedding approach combining both pretrained and morphologically-informed word embeddings performs best in the downstream task of Xhosa-English translation.

📄 PDF Abstract BibTeX arXiv:2003.04419

Code (0)

등록된 구현이 없습니다.

Tasks

TranslationWord Embeddings

Similar Papers 제목 키워드 기반

Sequence Tagging with Contextual and Non-Contextual Subword Representations: A Multilingual Evaluation

2019-06-04 · ACL 2019 7 · Benjamin Heinzerling, Michael Strube

Pretrained contextual and non-contextual subword embeddings have become available in over 250 languages, allowing massively multilingual NLP. However, while there is no dearth of pretrained embeddings, the distinct lack …

Multilingual Named Entity RecognitionMultilingual NLPnamed-entity-recognitionNamed Entity Recognition+2

Subword-based Cross-lingual Transfer of Embeddings from Hindi to Marathi

2021-10-16 · ACL ARR October 2021 10 · Anonymous

Word embeddings are growing to be a crucial resource in the field of NLP for any language. This work focuses on static subword embeddings transfer for Indian languages from a relatively higher resource language to a gene…

Cross-Lingual TransferWord EmbeddingsWord Similarity

LGSE: Lexically Grounded Subword Embedding Initialization for Low-Resource Language Adaptation

2026-03-23 · Hailay Teklehaymanot, Dren Fazlija, Wolfgang Nejdl arxiv

Adapting pretrained language models to low-resource, morphologically rich languages remains a significant challenge. Existing vocabulary expansion methods typically rely on arbitrarily segmented subword units, resulting …

Text ClassificationQuestion Answering

Subword-based Cross-lingual Transfer of Embeddings from Hindi to Marathi and Nepali

2022-07-01 · NAACL (SIGMORPHON) 2022 7 · Niyata Bafna, Zdeněk Žabokrtský

Word embeddings are growing to be a crucial resource in the field of NLP for any language. This work introduces a novel technique for static subword embeddings transfer for Indic languages from a relatively higher resour…

Cross-Lingual TransferWord EmbeddingsWord Similarity

WECHSEL: Effective initialization of subword embeddings for cross-lingual transfer of monolingual language models

2021-12-13 · NAACL 2022 7 · Benjamin Minixhofer, Fabian Paischer, Navid Rekabsaz

Large pretrained language models (LMs) have become the central building block of many NLP applications. Training these models requires ever more computational resources and most of the existing models are trained on Engl…

Cross-Lingual TransferWord Embeddings