Crowdsourcing and annotating NER for Twitter \#drift
We present two new NER datasets for Twitter; a manually annotated set of 1,467 tweets (kappa=0.942) and a set of 2,975 expert-corrected, crowdsourced NER annotated tweets from the dataset described in Finin et al. (2010). In our experiments with these datasets, we observe two important points: (a) language drift on Twitter is significant, and while off-the-shelf systems have been reported to perform well on in-sample data, they often perform poorly on new samples of tweets, (b) state-of-the-art performance across various datasets can be obtained from crowdsourced annotations, making it more feasible to {``}catch up{''} with language drift.
Code (0)
등록된 구현이 없습니다.
Tasks
Named Entity Recognition (NER)NERSimilar Papers 제목 키워드 기반
Annotating Targets of Opinions in Arabic using Crowdsourcing
A Crowdsourcing Approach for Annotating Causal Relation Instances in Wikipedia
CrowdMOT: Crowdsourcing Strategies for Tracking Multiple Objects in Videos
Crowdsourcing is a valuable approach for tracking objects in videos in a more scalable manner than possible with domain experts. However, existing frameworks do not produce high quality results with non-expert crowdworke…
DiversityAnnotating Sentiment and Irony in the Online Italian Political Debate on \#labuonascuola
In this paper we present the TWitterBuonaScuola corpus (TW-BS), a novel Italian linguistic resource for Sentiment Analysis, developed with the main aim of analyzing the online debate on the controversial Italian politica…
Sentiment Analysis