paper-with-me

Papers

Lexical Comparison Between Wikipedia and Twitter Corpora by Using Word Embeddings

2015-07-01 · IJCNLP 2015 7 · Luchen Tan, Haotian Zhang, Charles Clarke, Mark Smucker
📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Machine TranslationWord Embeddings

Similar Papers 제목 키워드 기반

Wikification for Scriptio Continua

2016-05-01 · LREC 2016 5 · Yugo Murawaki, Shinsuke Mori

The fact that Japanese employs scriptio continua, or a writing system without spaces, complicates the first step of an NLP pipeline. Word segmentation is widely used in Japanese language processing, and lexical knowledge…

Segmentation

VMWE discovery: a comparative analysis between Literature and Twitter Corpora

2020-12-01 · COLING (MWE) 2020 12 · Vivian Stamou, Artemis Xylogianni, Marilena Malli, Penny Takorou 외

We evaluate manually five lexical association measurements as regards the discovery of Modern Greek verb multiword expressions with two or more lexicalised components usingmwetoolkit3 (Ramisch et al., 2010). We use Twitt…

Using four different online media sources to forecast the crude oil price

2021-05-19 · M. Elshendy, A. Fronzetti Colladon, E. Battistoni, P. A. Gloor

This study looks for signals of economic awareness on online social media and tests their significance in economic predictions. The study analyses, over a period of two years, the relationship between the West Texas Inte…

Articles

Twitter as a Lifeline: Human-annotated Twitter Corpora for NLP of Crisis-related Messages

2016-05-19 · LREC 2016 5 · Muhammad Imran, Prasenjit Mitra, Carlos Castillo

Microblogging platforms such as Twitter provide active communication channels during mass convergence and emergency events such as earthquakes, typhoons. During the sudden onset of a crisis situation, affected people pos…

Disaster ResponseHumanitarianWord Embeddings

Cultural Topic Modelling over Novel Wikipedia Corpora for South-Slavic Languages

2021-09-01 · RANLP 2021 9 · Filip Markoski, Elena Markoska, Nikola Ljubešić, Eftim Zdravevski 외

There is a shortage of high-quality corpora for South-Slavic languages. Such corpora are useful to computer scientists and researchers in social sciences and humanities alike, focusing on numerous linguistic, content ana…

Cultural Vocal Bursts Intensity Prediction