paper-with-me

Papers

Benchmarking of Transformer-Based Pre-Trained Models on Social Media Text Classification Datasets

2020-12-01 · ALTA 2020 12 · Yuting Guo, Xiangjue Dong, Mohammed Ali Al-Garadi, Abeed Sarker, Cecile Paris, Diego Mollá Aliod

Free text data from social media is now widely used in natural language processing research, and one of the most common machine learning tasks performed on this data is classification. Generally speaking, performances of supervised classification algorithms on social media datasets are lower than those on texts from other sources, but recently-proposed transformer-based models have considerably improved upon legacy state-of-the-art systems. Currently, there is no study that compares the performances of different variants of transformer-based models on a wide range of social media text classification datasets. In this paper, we benchmark the performances of transformer-based pre-trained models on 25 social media text classification datasets, 6 of which are health-related. We compare three pre-trained language models, RoBERTa-base, BERTweet and ClinicalBioBERT in terms of classification accuracy. Our experiments show that RoBERTa-base and BERTweet perform comparably on most datasets, and considerably better than ClinicalBioBERT, even on health-related datasets.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

BenchmarkingClassificationtext-classificationText Classification

Similar Papers 제목 키워드 기반

Benchmarking for Public Health Surveillance tasks on Social Media with a Domain-Specific Pretrained Language Model

2022-04-09 · nlppower (ACL) 2022 5 · Usman Naseem, Byoung Chan Lee, Matloob Khushi, Jinman Kim 외

A user-generated text on social media enables health workers to keep track of information, identify possible outbreaks, forecast disease trends, monitor emergency cases, and ascertain disease awareness and response to of…

BenchmarkingLanguage ModelingLanguage Modelling

Benchmarking Language Models for Cyberbullying Identification and Classification from Social-media Texts

2022-06-01 · LATERAISSE (LREC) 2022 6 · Kanishk Verma, Tijana Milosevic, Keith Cortis, Brian Davis

Cyberbullying is bullying perpetrated via the medium of modern communication technologies like social media networks and gaming platforms. Unfortunately, most existing datasets focusing on cyberbullying detection or clas…

BenchmarkingBinary ClassificationClassification

MultiSocial: Multilingual Benchmark of Machine-Generated Text Detection of Social-Media Texts

2024-06-18 · Dominik Macko, Jakub Kopal, Robert Moro, Ivan Srba

Recent LLMs are able to generate high-quality multilingual texts, indistinguishable for humans from authentic human-written ones. Research in machine-generated text detection is however mostly focused on the English lang…

ArticlesBenchmarkingText Detection

ViSoBERT: A Pre-Trained Language Model for Vietnamese Social Media Text Processing

2023-10-17 · Quoc-Nam Nguyen, Thang Chau Phan, Duc-Vu Nguyen, Kiet Van Nguyen

English and Chinese, known as resource-rich languages, have witnessed the strong development of transformer-based language models for natural language processing tasks. Although Vietnam has approximately 100M people spea…

Language ModelingLanguage ModellingVietnamese Hate Speech DetectionVietnamese Language Models+2

CrisisBench: Benchmarking Crisis-related Social Media Datasets for Humanitarian Information Processing

2020-04-14 · Firoj Alam, Hassan Sajjad, Muhammad Imran, Ferda Ofli

Time-critical analysis of social media streams is important for humanitarian organizations for planing rapid response during disasters. The \textit{crisis informatics} research community has developed several techniques …

BenchmarkingGeneral ClassificationHumanitarianInformativeness