Time-Aware Word Embeddings for Three Lebanese News Archives
Word embeddings have proven to be an effective method for capturing semantic relations among distinct terms within a large corpus. In this paper, we present a set of word embeddings learnt from three large Lebanese news archives, which collectively consist of 609,386 scanned newspaper images and spanning a total of 151 years, ranging from 1933 till 2011. The diversified ideological nature of the news archives alongside the temporal variability of the embeddings offer a rare glimpse onto the variation of word representation across the left-right political spectrum. To train the word embeddings, Google{'}s Tesseract 4.0 OCR engine was employed to transcribe the scanned news archives, and various archive-level as well as decade-level word embeddings were learnt. To evaluate the accuracy of the learnt word embeddings, a benchmark of analogy tasks was used. Finally, we demonstrate an interactive system that allows the end user to visualize for a given word of interest, the variation of the top-k closest words in the embedding space as a function of time and across news archives using an animated scatter plot.
Code (1)
Tasks
Optical Character Recognition (OCR)Word EmbeddingsSimilar Papers 제목 키워드 기반
Fine-Tuning LLMs for Low-Resource Dialect Translation: The Case of Lebanese
This paper examines the effectiveness of Large Language Models (LLMs) in translating the low-resource Lebanese dialect, focusing on the impact of culturally authentic data versus larger translated datasets. We compare th…
TranslationDiaLex: A Benchmark for Evaluating Multidialectal Arabic Word Embeddings
Word embeddings are a core component of modern natural language processing systems, making the ability to thoroughly evaluate them a vital task. We describe DiaLex, a benchmark for intrinsic evaluation of dialectal Arabi…
Word EmbeddingsLexical Induction of Morphological and Orthographic Forms for Low-Resourced Languages
In this work we address the issue of high-degree lexical sparsity for non-standard languages under severe circumstance of small resources that are considered insufficient to train recent powerful language models. We prop…
Word EmbeddingsLEBANONUPRISING: a thorough study of Lebanese tweets
Recent studies showed a huge interest in social networks sentiment analysis. Twitter, which is a microblogging service, can be a great source of information on how the users feel about a certain topic, or what their opin…
Sentiment AnalysisLearning Domain-Sensitive and Sentiment-Aware Word Embeddings
Word embeddings have been widely used in sentiment classification because of their efficacy for semantic representations of words. Given reviews from different domains, some existing methods for word embeddings exploit s…
Data AugmentationGeneral ClassificationSentenceSentiment Analysis+2