paper-with-me

Broad Twitter Corpus

홈페이지 · 논문 12편

This paper introduces the Broad Twitter Corpus (BTC), which is not only significantly bigger, but sampled across different regions, temporal periods, and types of Twitter users. The gold-standard named entity annotations are made by a combination of NLP experts and crowd workers, which enables us to harness crowd recall while maintaining high quality. We also measure the entity drift observed in our dataset (i.e. how entity representation varies over time), and compare to newswire.

Texts English

벤치마크

Zero-shot Named Entity Recognition (NER) on Broad Twitter Corpus 결과 2개
Named Entity Recognition (NER) on Broad Twitter Corpus 결과 1개