Turkish Tweet Classification with Transformer Encoder
Short-text classification is a challenging task, due to the sparsity and high dimensionality of the feature space. In this study, we aim to analyze and classify Turkish tweets based on their topics. Social media jargon and the agglutinative structure of the Turkish language makes this classification task even harder. As far as we know, this is the first study that uses a Transformer Encoder for short text classification in Turkish. The model is trained in a weakly supervised manner, where the training data set has been labeled automatically. Our results on the test set, which has been manually labeled, show that performing morphological analysis improves the classification performance of the traditional machine learning algorithms Random Forest, Naive Bayes, and Support Vector Machines. Still, the proposed approach achieves an F-score of 89.3 {\%} outperforming those algorithms by at least 5 points.
Code (0)
등록된 구현이 없습니다.
Tasks
ClassificationGeneral ClassificationMorphological Analysistext-classificationText ClassificationMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
TurkishBERTweet: Fast and Reliable Large Language Model for Social Media Analysis
Turkish is one of the most popular languages in the world. Wide us of this language on social media platforms such as Twitter, Instagram, or Tiktok and strategic position of the country in the world politics makes it app…
Hate Speech DetectionLanguage ModelingLanguage ModellingLarge Language Model+4A Turkish Hate Speech Dataset and Detection System
Social media posts containing hate speech are reproduced and redistributed at an accelerated pace, reaching greater audiences at a higher speed. We present a machine learning system for automatic detection of hate speech…
Binary ClassificationHate Speech DetectionTabiBERT: A Large-Scale ModernBERT Foundation Model and A Unified Benchmark for Turkish
Since the inception of BERT, encoder-only Transformers have evolved significantly in computational efficiency, training stability, and long-context modeling. ModernBERT consolidates these advances by integrating Rotary P…
Computational EfficiencyDomain GeneralizationQuestion AnsweringFine-tuning Transformer-based Encoder for Turkish Language Understanding Tasks
Deep learning-based and lately Transformer-based language models have been dominating the studies of natural language processing in the last years. Thanks to their accurate and fast fine-tuning characteristics, they have…
named-entity-recognitionNamed Entity RecognitionNatural Language UnderstandingQuestion Answering+4CoLi at UdS at SemEval-2020 Task 12: Offensive Tweet Detection with Ensembling
We present our submission and results for SemEval-2020 Task 12: Multilingual Offensive Language Identification in Social Media (OffensEval 2020) where we participated in offensive tweet classification tasks in English, A…
BIG-bench Machine LearningLanguage IdentificationregressionXLM-R