paper-with-me

Papers

A Fast, Compact, Accurate Model for Language Identification of Codemixed Text

2018-10-09 · EMNLP 2018 10 · Yuan Zhang, Jason Riesa, Daniel Gillick, Anton Bakalov, Jason Baldridge, David Weiss

We address fine-grained multilingual language identification: providing a language code for every token in a sentence, including codemixed text containing multiple languages. Such text is prevalent online, in documents, social media, and message boards. We show that a feed-forward network with a simple globally constrained decoder can accurately and rapidly label both codemixed and monolingual text in 100 languages and 100 language pairs. This model outperforms previously published multilingual approaches in terms of both accuracy and speed, yielding an 800x speed-up and a 19.5% averaged absolute gain on three codemixed datasets. It furthermore outperforms several benchmark systems on monolingual language identification.

📄 PDF Abstract BibTeX arXiv:1810.04142

Code (0)

등록된 구현이 없습니다.

Tasks

DecoderLanguage IdentificationSentence

Similar Papers 제목 키워드 기반

Language Identification of Hindi-English tweets using code-mixed BERT

2021-07-02 · Mohd Zeeshan Ansari, M M Sufyan Beg, Tanvir Ahmad, Mohd Jazib Khan 외

Language identification of social media text has been an interesting problem of study in recent years. Social media messages are predominantly in code mixed in non-English speaking states. Prior knowledge by pre-training…

Language IdentificationTransfer Learning

Small and Practical BERT Models for Sequence Labeling

2019-08-31 · IJCNLP 2019 11 · Henry Tsai, Jason Riesa, Melvin Johnson, Naveen Arivazhagan 외

We propose a practical scheme to train a single multilingual sequence labeling model that yields state of the art results and is small and fast enough to run on a single CPU. Starting from a public multilingual BERT chec…

CPUPart-Of-Speech Tagging

KanCMD: Kannada CodeMixed Dataset for Sentiment Analysis and Offensive Language Detection

2020-12-01 · COLING (PEOPLES) 2020 12 · Adeep Hande, Ruba Priyadharshini, Bharathi Raja Chakravarthi

We introduce Kannada CodeMixed Dataset (KanCMD), a multi-task learning dataset for sentiment analysis and offensive language identification. The KanCMD dataset highlights two real-world issues from the social media text.…

Language IdentificationMulti-Task LearningSentiment Analysis

Findings of the Shared Task on Offensive Span Identification from Code-Mixed Tamil-English Comments

2022-05-12 · Manikandan Ravikiran, Bharathi Raja Chakravarthi, Anand Kumar Madasamy, Sangeetha Sivanesan 외

Offensive content moderation is vital in social media platforms to support healthy online discussions. However, their prevalence in codemixed Dravidian languages is limited to classifying whole comments without identifyi…

Curriculum Learning Strategies for Hindi-English Codemixed Sentiment Analysis

2019-06-18 · Anirudh Dahiya, Neeraj Battan, Manish Shrivastava, Dipti Mishra Sharma

Sentiment Analysis and other semantic tasks are commonly used for social media textual analysis to gauge public opinion and make sense from the noise on social media. The language used on social media not only commonly d…

Sentiment Analysis