Harvey Mudd College at SemEval-2019 Task 4: The D.X. Beaumont Hyperpartisan News Detector
We use the 600 hand-labelled articles from SemEval Task 4 to hand-tune a classifier with 3000 features for the Hyperpartisan News Detection task. Our final system uses features based on bag-of-words (BoW), analysis of the article title, language complexity, and simple sentiment analysis in a naive Bayes classifier. We trained our final system on the 600,000 articles labelled by publisher. Our final system has an accuracy of 0.653 on the hand-labeled test set. The most effective features are the Automated Readability Index and the presence of certain words in the title. This suggests that hyperpartisan writing uses a distinct writing style, especially in the title.
Code (0)
등록된 구현이 없습니다.
Tasks
ArticlesSentiment AnalysisSimilar Papers 제목 키워드 기반
Harvey Mudd College at SemEval-2019 Task 4: The Carl Kolchak Hyperpartisan News Detector
We use various natural processing and machine learning methods to perform the Hyperpartisan News Detection task. In particular, some of the features we look at are bag-of-words features, the title{'}s length, number of c…
Articlesfeature selectionSentiment AnalysisHarvey Mudd College at SemEval-2019 Task 4: The Clint Buchanan Hyperpartisan News Detector
We investigate the recently developed Bidirectional Encoder Representations from Transformers (BERT) model for the hyperpartisan news detection task. Using a subset of hand-labeled articles from SemEval as a validation s…
ArticlesJCTICOL at SemEval-2019 Task 6: Classifying Offensive Language in Social Media using Deep Learning Methods, Word/Character N-gram Features, and Preprocessing Methods
In this paper, we describe our submissions to SemEval-2019 task 6 contest. We tackled all three sub-tasks in this task {``}OffensEval - Identifying and Categorizing Offensive Language in Social Media{''}. In our system c…
PositionMUDD: A New Re-Identification Dataset with Efficient Annotation for Off-Road Racers in Extreme Conditions
Re-identifying individuals in unconstrained environments remains an open challenge in computer vision. We introduce the Muddy Racer re-IDentification Dataset (MUDD), the first large-scale benchmark for matching identitie…
Sports AnalyticsJCTDHS at SemEval-2019 Task 5: Detection of Hate Speech in Tweets using Deep Learning Methods, Character N-gram Features, and Preprocessing Methods
In this paper, we describe our submissions to SemEval-2019 contest. We tackled subtask A - {``}a binary classification where systems have to predict whether a tweet with a given target (women or immigrants) is hateful or…
Binary ClassificationGeneral ClassificationPosition