Papers Native Language Identification
“Native Language Identification” 태그가 달린 논문 88편 · 필터 해제
On the Development of a Large Scale Corpus for Native Language Identification
Native Language Identification (NLI) is the task of identifying an author’s native language from their writings in a second language. In this paper, we introduce a new corpus (italki), which is larger than the current co…
BIG-bench Machine LearningLanguage IdentificationNative Language IdentificationTransductive Learning with String Kernels for Cross-Domain Text Classification
For many text classification tasks, there is a major problem posed by the lack of labeled data in a target domain. Although classifiers for a target domain can be trained on labeled text data from a related source domain…
ClassificationCross-Domain Text ClassificationGeneral ClassificationLanguage Identification+4Native Language Identification with User Generated Content
We address the task of native language identification in the context of social media content, where authors are highly-fluent, advanced nonnative speakers (of English). Using both linguistically-motivated features and th…
Language IdentificationNative Language IdentificationThe Role of Emotions in Native Language Identification
We explore the hypothesis that emotion is one of the dimensions of language that surfaces from the native language into a second language. To check the role of emotions in native language identification (NLI), we model e…
Deception DetectionLanguage IdentificationNative Language IdentificationSentiment AnalysisNative Language Identification With Classifier Stacking and Ensembles
Ensemble methods using multiple classifiers have proven to be among the most successful approaches for the task of Native Language Identification (NLI), achieving the current state of the art. However, a systematic exami…
Cross-corpusGeneral ClassificationLanguage AcquisitionLanguage Identification+2Improving the results of string kernels in sentiment analysis and Arabic dialect identification by adapting them to your test set
Recently, string kernels have obtained state-of-the-art results in various text classification tasks such as Arabic dialect identification or native language identification. In this paper, we apply two simple yet effecti…
Dialect IdentificationGeneral ClassificationLanguage IdentificationNative Language Identification+4Punctuation as Native Language Interference
In this paper, we describe experiments designed to explore and evaluate the impact of punctuation marks on the task of native language identification. Punctuation is specific to each language, and is part of the indicato…
ClassificationCross-corpusGeneral ClassificationLanguage Identification+2Predicting Foreign Language Usage from English-Only Social Media Posts
Social media is known for its multi-cultural and multilingual interactions, a natural product of which is code-mixing. Multilingual speakers mix languages they tweet to address a different audience, express certain feeli…
Cross-Lingual TransferLanguage IdentificationNative Language IdentificationTransfer LearningUsing Classifier Features to Determine Language Transfer on Morphemes
The aim of this thesis is to perform a Native Language Identification (NLI) task where we identify an English learner{'}s native language background based only on the learner{'}s English writing samples. We focus on the …
Cross-corpusLanguage AcquisitionLanguage IdentificationNative Language Identification+1Cross-corpus Native Language Identification via Statistical Embedding
In this paper, we approach the task of native language identification in a realistic cross-corpus scenario where a model is trained with available data and has to predict the native language from data of a different corp…
Cross-corpusLanguage IdentificationNative Language IdentificationA Portuguese Native Language Identification Dataset
In this paper we present NLI-PT, the first Portuguese dataset compiled for Native Language Identification (NLI), the task of identifying an author's first language based on their second language writing. The dataset incl…
Language AcquisitionLanguage IdentificationNative Language IdentificationPOSAutomated essay scoring with string kernels and word embeddings
In this work, we present an approach based on combining string kernels and word embeddings for automatic essay scoring. String kernels capture the similarity among strings based on counting common character n-grams, whic…
Automated Essay ScoringDialect IdentificationGeneral ClassificationLanguage Identification+4A Report on the 2017 Native Language Identification Shared Task
Native Language Identification (NLI) is the task of automatically identifying the native language (L1) of an individual based on their language production in a learned language. It is typically framed as a classification…
Grammatical Error CorrectionLanguage AcquisitionLanguage IdentificationNative Language IdentificationNative Language Identification Using a Mixture of Character and Word N-grams
Native language identification (NLI) is the task of determining an author{'}s native language, based on a piece of his/her writing in a second language. In recent years, NLI has received much attention due to its challen…
Language AcquisitionLanguage IdentificationNative Language IdentificationEnsemble Methods for Native Language Identification
Our team{---}Uvic-NLP{---}explored and evaluated a variety of lexical features for Native Language Identification (NLI) within the framework of ensemble methods. Using a subset of the highest performing features, we trai…
Language AcquisitionLanguage IdentificationNative Language IdentificationNeural Networks and Spelling Features for Native Language Identification
We present the RUG-SU team{'}s submission at the Native Language Identification Shared Task 2017. We combine several approaches into an ensemble, based on spelling error features, a simple neural network using word repre…
Language IdentificationNative Language IdentificationWord EmbeddingsA study of N-gram and Embedding Representations for Native Language Identification
We report on our experiments with N-gram and embedding based feature representations for Native Language Identification (NLI) as a part of the NLI Shared Task 2017 (team name: NLI-ISU). Our best performing system on the …
Feature EngineeringLanguage AcquisitionLanguage IdentificationNative Language Identification+1A Shallow Neural Network for Native Language Identification with Character N-grams
This paper describes the systems submitted by GadjahMada team to the Native Language Identification (NLI) Shared Task 2017. Our models used a continuous representation of character n-grams which are learned jointly with …
Language IdentificationNative Language IdentificationFewer features perform well at Native Language Identification task
This paper describes our results at the NLI shared task 2017. We participated in essays, speech, and fusion task that uses text, speech, and i-vectors for the task of identifying the native language of the given input. I…
Language IdentificationNative Language IdentificationExploring Optimal Voting in Native Language Identification
We describe the submissions entered by the National Research Council Canada in the NLI-2017 evaluation. We mainly explored the use of voting, and various ways to optimize the choice and number of voting systems. We also …
Language IdentificationNative Language Identification