A Fast, Compact, Accurate Model for Language Identification of Codemixed Text
We address fine-grained multilingual language identification: providing a language code for every token in a sentence, including codemixed text containing multiple languages. Such text is prevalent online, in documents, social media, and message boards. We show that a feed-forward network with a simple globally constrained decoder can accurately and rapidly label both codemixed and monolingual text in 100 languages and 100 language pairs. This model outperforms previously published multilingual approaches in terms of both accuracy and speed, yielding an 800x speed-up and a 19.5% averaged absolute gain on three codemixed datasets. It furthermore outperforms several benchmark systems on monolingual language identification.
Code (0)
등록된 구현이 없습니다.
Tasks
DecoderLanguage IdentificationSentenceSimilar Papers 제목 키워드 기반
Language Identification of Hindi-English tweets using code-mixed BERT
Language identification of social media text has been an interesting problem of study in recent years. Social media messages are predominantly in code mixed in non-English speaking states. Prior knowledge by pre-training…
Language IdentificationTransfer LearningSmall and Practical BERT Models for Sequence Labeling
We propose a practical scheme to train a single multilingual sequence labeling model that yields state of the art results and is small and fast enough to run on a single CPU. Starting from a public multilingual BERT chec…
CPUPart-Of-Speech TaggingKanCMD: Kannada CodeMixed Dataset for Sentiment Analysis and Offensive Language Detection
We introduce Kannada CodeMixed Dataset (KanCMD), a multi-task learning dataset for sentiment analysis and offensive language identification. The KanCMD dataset highlights two real-world issues from the social media text.…
Language IdentificationMulti-Task LearningSentiment AnalysisFindings of the Shared Task on Offensive Span Identification from Code-Mixed Tamil-English Comments
Offensive content moderation is vital in social media platforms to support healthy online discussions. However, their prevalence in codemixed Dravidian languages is limited to classifying whole comments without identifyi…
Curriculum Learning Strategies for Hindi-English Codemixed Sentiment Analysis
Sentiment Analysis and other semantic tasks are commonly used for social media textual analysis to gauge public opinion and make sense from the noise on social media. The language used on social media not only commonly d…
Sentiment Analysis