paper-with-me

홈 › Papers

An Embarrassingly Simple Method to Mitigate Undesirable Properties of Pretrained Language Model Tokenizers

2022-05-01 · ACL 2022 5 · Valentin Hofmann, Hinrich Schuetze, Janet Pierrehumbert

We introduce FLOTA (Few Longest Token Approximation), a simple yet effective method to improve the tokenization of pretrained language models (PLMs). FLOTA uses the vocabulary of a standard tokenizer but tries to preserve the morphological structure of words during tokenization. We evaluate FLOTA on morphological gold segmentations as well as a text classification task, using BERT, GPT-2, and XLNet as example PLMs. FLOTA leads to performance gains, makes inference more efficient, and enhances the robustness of PLMs with respect to whitespace noise.

📄 PDF Abstract BibTeX

Code (1)

valentinhofmann/flota 공식 구현 pytorch

Tasks

Language ModelingLanguage Modellingtext-classificationText Classification

Similar Papers 제목 키워드 기반

Sliced-Wasserstein Autoencoder: An Embarrassingly Simple Generative Model

2018-04-05 · Soheil Kolouri, Phillip E. Pope, Charles E. Martin, Gustavo K. Rohde

In this paper we study generative modeling via autoencoders while using the elegant geometric properties of the optimal transport (OT) problem and the Wasserstein distances. We introduce Sliced-Wasserstein Autoencoders (…

model

Sliced-Wasserstein Autoencoder: An Embarrassingly Simple Generative Model

2018-04-05 · Anonymous

In this paper we study generative modeling via autoencoders while using the elegant geometric properties of the optimal transport (OT) problem and the Wasserstein distances. We introduce Sliced-Wasserstein Autoencoders (…

model

Learning from the Undesirable: Robust Adaptation of Language Models without Forgetting

2025-11-17 · Yunhun Nam, Jaehyung Kim, Jongheon Jeong arxiv

Language models (LMs) are often adapted through supervised fine-tuning (SFT) to specialize their capabilities for downstream tasks. However, in typical scenarios where the fine-tuning data is limited, e.g., compared to p…

Data Augmentation

An Embarrassingly Simple Approach for Transfer Learning from Pretrained Language Models

2019-02-27 · NAACL 2019 6 · Alexandra Chronopoulou, Christos Baziotis, Alexandros Potamianos

A growing number of state-of-the-art transfer learning methods employ language models pretrained on large generic corpora. In this paper we present a conceptually simple and effective transfer learning approach that addr…

General ClassificationLanguage ModelingLanguage Modellingtext-classification+2

Embarrassingly Easy Document-Level MT Metrics: How to Convert Any Pretrained Metric Into a Document-Level Metric

2022-09-27 · Giorgos Vernikos, Brian Thompson, Prashant Mathur, Marcello Federico

We hypothesize that existing sentence-level machine translation (MT) metrics become less effective when the human reference contains ambiguities. To verify this hypothesis, we present a very simple method for extending p…

Machine TranslationSentence