paper-with-me

홈 › Papers

Improving Character-Aware Neural Language Model by Warming up Character Encoder under Skip-gram Architecture

2021-09-01 · RANLP 2021 9 · Yukun Feng, Chenlong Hu, Hidetaka Kamigaito, Hiroya Takamura, Manabu Okumura

Character-aware neural language models can capture the relationship between words by exploiting character-level information and are particularly effective for languages with rich morphology. However, these models are usually biased towards information from surface forms. To alleviate this problem, we propose a simple and effective method to improve a character-aware neural language model by forcing a character encoder to produce word-based embeddings under Skip-gram architecture in a warm-up step without extra training data. We empirically show that the resulting character-aware neural language model achieves obvious improvements of perplexity scores on typologically diverse languages, that contain many low-frequency or unseen words.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Language ModelingLanguage Modelling

Similar Papers 제목 키워드 기반

WarmServe: Enabling One-for-Many GPU Prewarming for Multi-LLM Serving

2025-12-10 · Chiheng Lou, Sheng Qi, Rui Kang, Yong Zhang 외 arxiv

Deploying multiple models within shared GPU clusters is a key strategy to improve resource efficiency in large language model (LLM) serving. Existing multi-LLM serving systems improve GPU utilization at the cost of degra…

FakeSwarm: Improving Fake News Detection with Swarming Characteristics

2023-05-30 · Jun Wu, Xuesong Ye

The proliferation of fake news poses a serious threat to society, as it can misinform and manipulate the public, erode trust in institutions, and undermine democratic processes. To address this issue, we present FakeSwar…

Fake News Detection

A Character-Aware Encoder for Neural Machine Translation

2016-12-01 · COLING 2016 12 · Zhen Yang, Wei Chen, Feng Wang, Bo Xu

This article proposes a novel character-aware neural machine translation (NMT) model that views the input sequences as sequences of characters rather than words. On the use of row convolution (Amodei et al., 2015), the e…

Machine TranslationNMTTranslation

Swarming for Faster Convergence in Stochastic Optimization

2018-06-11 · Shi Pu, Alfredo Garcia

We study a distributed framework for stochastic optimization which is inspired by models of collective motion found in nature (e.g., swarming) with mild communication requirements. Specifically, we analyze a scheme in wh…

Stochastic Optimization

Text Classification through Glyph-aware Disentangled Character Embedding and Semantic Sub-character Augmentation

2020-11-09 · Asian Chapter of the Association for Computational Linguistics 2020 · Takumi Aoki, Shunsuke Kitada, Hitoshi Iyatomi

We propose a new character-based text classification framework for non-alphabetic languages, such as Chinese and Japanese. Our framework consists of a variational character encoder (VCE) and character-level text classifi…

ClassificationData AugmentationGeneral ClassificationSentence+2