Neural Temporality Adaptation for Document Classification: Diachronic Word Embeddings and Domain Adaptation Models
Language usage can change across periods of time, but document classifiers models are usually trained and tested on corpora spanning multiple years without considering temporal variations. This paper describes two complementary ways to adapt classifiers to shifts across time. First, we show that diachronic word embeddings, which were originally developed to study language change, can also improve document classification, and we show a simple method for constructing this type of embedding. Second, we propose a time-driven neural classification model inspired by methods for domain adaptation. Experiments on six corpora show how these methods can make classifiers more robust over time.
Code (1)
Tasks
ClassificationDiachronic Word EmbeddingsDocument ClassificationDomain AdaptationGeneral ClassificationWord EmbeddingsSimilar Papers 제목 키워드 기반
Temporal Adaptation of BERT and Performance on Downstream Document Classification: Insights from Social Media
Language use differs between domains and even within a domain, language use changes over time. For pre-trained language models like BERT, domain adaptation through continued pre-training has been shown to improve perform…
Document ClassificationDomain AdaptationGeneral ClassificationLanguage ModellingHow Diachronic Text Corpora Affect Context based Retrieval of OOV Proper Names for Audio News
Out-Of-Vocabulary (OOV) words missed by Large Vocabulary Continuous Speech Recognition (LVCSR) systems can be recovered with the help of topic and semantic context of the OOV words captured from a diachronic text corpus.…
Retrievalspeech-recognitionSpeech RecognitionExamining Temporality in Document Classification
Many corpora span broad periods of time. Language processing models trained during one time period may not work well in future time periods, and the best model may depend on specific times of year (e.g., people might des…
ClassificationDocument ClassificationDomain AdaptationGeneral ClassificationDHPLT: large-scale multilingual diachronic corpora and word representations for semantic change modelling
In this resource paper, we present DHPLT, an open collection of diachronic corpora in 41 diverse languages. DHPLT is based on the web-crawled HPLT datasets; we use web crawl timestamps as the approximate signal of docume…
Roadblocks in Gender Bias Measurement for Diachronic Corpora
The use of word embeddings is an important NLP technique for extracting meaningful conclusions from corpora of human text. One important question that has been raised about word embeddings is the degree of gender bias le…
Word Embeddings