Predicting Foreign Language Usage from English-Only Social Media Posts
Social media is known for its multi-cultural and multilingual interactions, a natural product of which is code-mixing. Multilingual speakers mix languages they tweet to address a different audience, express certain feelings, or attract attention. This paper presents a large-scale analysis of 6 million tweets produced by 27 thousand multilingual users speaking 12 other languages besides English. We rely on this corpus to build predictive models to infer non-English languages that users speak exclusively from their English tweets. Unlike native language identification task, we rely on large amounts of informal social media communications rather than ESL essays. We contrast the predictive power of the state-of-the-art machine learning models trained on lexical, syntactic, and stylistic signals with neural network models learned from word, character and byte representations extracted from English only tweets. We report that content, style and syntax are the most predictive of non-English languages that users speak on Twitter. Neural network models learned from byte representations of user content combined with transfer learning yield the best performance. Finally, by analyzing cross-lingual transfer {--} the influence of non-English languages on various levels of linguistic performance in English, we present novel findings on stylistic and syntactic variations across speakers of 12 languages in social media.
Code (0)
등록된 구현이 없습니다.
Tasks
Cross-Lingual TransferLanguage IdentificationNative Language IdentificationTransfer LearningSimilar Papers 제목 키워드 기반
Machine Translation Of Bi-Lingual Hindi-English (Hinglish) text. 10th Machine Translation summit (MT Summit X),
In the present communication-based society, no natural language seems to have been left untouched by the trends of code-mixing. For different communicative purposes, a language uses linguistic codes from other langua…
Machine TranslationSentenceTranslationA Chinese Writing Correction System for Learning Chinese as a Foreign Language
We present a Chinese writing correction system for learning Chinese as a foreign language. The system takes a wrong input sentence and generates several correction suggestions. It also retrieves example Chinese sentences…
Grammatical Error CorrectionMachine TranslationSentenceTranslation+1Problems with the use of Web search engines to find results in foreign languages
Purpose - To test the ability of major search engines, Google, Yahoo, MSN, and Ask, to distinguish between German and English-language documents Design/methodology/approach - 50 queries, using words common in German an…
Pronunciation Modeling of Foreign Words for Mandarin ASR by Considering the Effect of Language Transfer
One of the challenges in automatic speech recognition is foreign words recognition. It is observed that a speaker's pronunciation of a foreign word is influenced by his native language knowledge, and such phenomenon is k…
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech RecognitionReconstructing Native Language Typology from Foreign Language Usage
Linguists and psychologists have long been studying cross-linguistic transfer, the influence of native language properties on linguistic performance in a foreign language. In this work we provide empirical evidence for t…