paper-with-me

Papers

Location-based Twitter Filtering for the Creation of Low-Resource Language Datasets in Indonesian Local Languages

2022-06-15 · Mukhlis Amien, Chong Feng, Heyan Huang

Twitter contains an abundance of linguistic data from the real world. We examine Twitter for user-generated content in low-resource languages such as local Indonesian. For NLP to work in Indonesian, it must consider local dialects, geographic context, and regional culture influence Indonesian languages. This paper identifies the problems we faced when constructing a Local Indonesian NLP dataset. Furthermore, we are developing a framework for creating, collecting, and classifying Local Indonesian datasets for NLP. Using twitter's geolocation tool for automatic annotating.

📄 PDF Abstract BibTeX arXiv:2206.07238

Code (0)

등록된 구현이 없습니다.

Tasks

Cultural Vocal Bursts Intensity Prediction

Similar Papers 제목 키워드 기반

Changes in Tweet Geolocation over Time: A Study with Carmen 2.0

2022-10-01 · COLING (WNUT) 2022 10 · Jingyu Zhang, Alexandra DeLucia, Mark Dredze

Researchers across disciplines use Twitter geolocation tools to filter data for desired locations. These tools have largely been trained and tested on English tweets, often originating in the United States from almost a …

Twitter corpus of Resource-Scarce Languages for Sentiment Analysis and Multilingual Emoji Prediction

2018-08-01 · COLING 2018 8 · Nurendra Choudhary, Rajat Singh, Vijjini Anvesh Rao, Manish Shrivastava

In this paper, we leverage social media platforms such as twitter for developing corpus across multiple languages. The corpus creation methodology is applicable for resource-scarce languages provided the speakers of that…

Sentiment Analysis

A New Twitter Verb Lexicon for Natural Language Processing

2012-05-01 · LREC 2012 5 · Jennifer Williams, Graham Katz

We describe in-progress work on the creation of a new lexical resource that contains a list of 486 verbs annotated with quantified temporal durations for the events that they describe. This resource is being compiled fro…

Game of Chess

Monolingual corpus creation and evaluation of truly low-resource languages from Peru

2020-07-01 · WS 2020 7 · Gina Bustamante, Arturo Oncevay

We introduce new monolingual corpora for four indigenous and endangered languages from Peru: Shipibo-konibo, Ashaninka, Yanesha and Yine. Given the total absence of these languages in the web, the extraction and processi…

Language Modelling

Creation of Corpus and analysis in Code-Mixed Kannada-English Twitter data for Emotion Prediction

2020-12-01 · COLING 2020 8 · Abhinav Reddy Appidi, Vamshi Krishna Srirangam, Darsi Suhas, Manish Shrivastava

Emotion prediction is a critical task in the field of Natural Language Processing (NLP). There has been a significant amount of work done in emotion prediction for resource-rich languages. There has been work done on cod…

Prediction