paper-with-me

Papers

Improving Classifier Training Efficiency for Automatic Cyberbullying Detection with Feature Density

2021-11-02 · Juuso Eronen, Michal Ptaszynski, Fumito Masui, Aleksander Smywiński-Pohl, Gniewosz Leliwa, Michal Wroczynski

We study the effectiveness of Feature Density (FD) using different linguistically-backed feature preprocessing methods in order to estimate dataset complexity, which in turn is used to comparatively estimate the potential performance of machine learning (ML) classifiers prior to any training. We hypothesise that estimating dataset complexity allows for the reduction of the number of required experiments iterations. This way we can optimize the resource-intensive training of ML models which is becoming a serious issue due to the increases in available dataset sizes and the ever rising popularity of models based on Deep Neural Networks (DNN). The problem of constantly increasing needs for more powerful computational resources is also affecting the environment due to alarmingly-growing amount of CO2 emissions caused by training of large-scale ML models. The research was conducted on multiple datasets, including popular datasets, such as Yelp business review dataset used for training typical sentiment analysis models, as well as more recent datasets trying to tackle the problem of cyberbullying, which, being a serious social problem, is also a much more sophisticated problem form the point of view of linguistic representation. We use cyberbullying datasets collected for multiple languages, namely English, Japanese and Polish. The difference in linguistic complexity of datasets allows us to additionally discuss the efficacy of linguistically-backed word preprocessing.

📄 PDF Abstract BibTeX arXiv:2111.01689

Code (0)

등록된 구현이 없습니다.

Tasks

Sentiment Analysis

Similar Papers 제목 키워드 기반

Automatic Detection of Cyberbullying in Social Media Text

2018-01-17 · Cynthia Van Hee, Gilles Jacobs, Chris Emmery, Bart Desmet 외

While social media offer great communication opportunities, they also increase the vulnerability of young people to threatening situations online. Recent studies report that cyberbullying constitutes a growing problem am…

Binary Classification

Initial Study into Application of Feature Density and Linguistically-backed Embedding to Improve Machine Learning-based Cyberbullying Detection

2022-06-04 · Juuso Eronen, Michal Ptaszynski, Fumito Masui, Gniewosz Leliwa 외

In this research, we study the change in the performance of machine learning (ML) classifiers when various linguistic preprocessing methods of a dataset were used, with the specific focus on linguistically-backed embeddi…

Aggressive, Repetitive, Intentional, Visible, and Imbalanced: Refining Representations for Cyberbullying Classification

2020-04-04 · Caleb Ziems, Ymir Vigfusson, Fred Morstatter

Cyberbullying is a pervasive problem in online communities. To identify cyberbullying cases in large-scale social networks, content moderators depend on machine learning classifiers for automatic cyberbullying detection.…

General Classification

Data Expansion Using WordNet-based Semantic Expansion and Word Disambiguation for Cyberbullying Detection

2022-06-01 · LREC 2022 6 · Md Saroar Jahan, Djamila Romaissa Beddiar, Mourad Oussalah, Muhidin Mohamed

Automatic identification of cyberbullying from textual content is known to be a challenging task. The challenges arise from the inherent structure of cyberbullying and the lack of labeled large-scale corpus, enabling eff…

Binary ClassificationData AugmentationWord Sense Disambiguation

Synthetic vs. Gold: The Role of LLM-Generated Labels and Data in Cyberbullying Detection

2025-02-21 · Arefeh Kazemi, Sri Balaaji Natarajan Kalaivendan, Joachim Wagner, Hamza Qadeer 외

This study investigates the role of LLM-generated synthetic data in cyberbullying detection. We conduct a series of experiments where we replace some or all of the authentic data with synthetic data, or augment the authe…