paper-with-me

홈 › Papers

Making the Most of Tweet-Inherent Features for Social Spam Detection on Twitter

2015-03-25 · Wang Bo, Zubiaga Arkaitz, Liakata Maria, Procter Rob

Social spam produces a great amount of noise on social media services such as Twitter, which reduces the signal-to-noise ratio that both end users and data mining applications observe. Existing techniques on social spam detection have focused primarily on the identification of spam accounts by using extensive historical and network-based data. In this paper we focus on the detection of spam tweets, which optimises the amount of data that needs to be gathered by relying only on tweet-inherent features. This enables the application of the spam detection system to a large set of tweets in a timely fashion, potentially applicable in a real-time or near real-time setting. Using two large hand-labelled datasets of tweets containing spam, we study the suitability of five classification algorithms and four different feature sets to the social spam detection task. Our results show that, by using the limited set of features readily available in a tweet, we can achieve encouraging results which are competitive when compared against existing spammer detection systems that make use of additional, costly user features. Our study is the first that attempts at generalising conclusions on the optimal classifiers and sets of features for social spam detection over different datasets.

📄 PDF Abstract BibTeX arXiv:1503.07405

Code (0)

등록된 구현이 없습니다.

Tasks

Spam detection

Similar Papers 제목 키워드 기반

Towards Real-Time, Country-Level Location Classification of Worldwide Tweets

2016-04-25 · Arkaitz Zubiaga, Alex Voss, Rob Procter, Maria Liakata 외

In contrast to much previous work that has focused on location classification of tweets restricted to a specific country, here we undertake the task in a broader context by classifying global tweets at the country level,…

ClassificationGeneral Classification

Discourse-Aware Rumour Stance Classification in Social Media Using Sequential Classifiers

2017-12-06 · Arkaitz Zubiaga, Elena Kochkina, Maria Liakata, Rob Procter 외

Rumour stance classification, defined as classifying the stance of specific social media posts into one of supporting, denying, querying or commenting on an earlier post, is becoming of increasing interest to researchers…

General ClassificationStance Classification

TurkishBERTweet: Fast and Reliable Large Language Model for Social Media Analysis

2023-11-29 · Ali Najafi, Onur Varol

Turkish is one of the most popular languages in the world. Wide us of this language on social media platforms such as Twitter, Instagram, or Tiktok and strategic position of the country in the world politics makes it app…

Hate Speech DetectionLanguage ModelingLanguage ModellingLarge Language Model+4

Understanding Information Spreading Mechanisms During COVID-19 Pandemic by Analyzing the Impact of Tweet Text and User Features for Retweet Prediction

2021-05-26 · Pervaiz Iqbal Khan, Imran Razzak, Andreas Dengel, Sheraz Ahmed

COVID-19 has affected the world economy and the daily life routine of almost everyone. It has been a hot topic on social media platforms such as Twitter, Facebook, etc. These social media platforms enable users to share …

Stylometric Detection of AI-Generated Text in Twitter Timelines

2023-03-07 · Tharindu Kumarage, Joshua Garland, Amrita Bhattacharjee, Kirill Trapeznikov 외

Recent advancements in pre-trained language models have enabled convenient methods for generating human-like text at a large scale. Though these generation capabilities hold great potential for breakthrough applications,…

Language ModellingMisinformationTask 2