Papers Native Language Identification
“Native Language Identification” 태그가 달린 논문 88편 · 필터 해제
CIC-FBK Approach to Native Language Identification
We present the CIC-FBK system, which took part in the Native Language Identification (NLI) Shared Task 2017. Our approach combines features commonly used in previous NLI research, i.e., word n-grams, lemma n-grams, part-…
General ClassificationLanguage IdentificationLEMMANative Language IdentificationThe Power of Character N-grams in Native Language Identification
In this paper, we explore the performance of a linear SVM trained on language independent character features for the NLI Shared Task 2017. Our basic system (GRONINGEN) achieves the best performance (87.56 F1-score) on th…
Language IdentificationNative Language IdentificationText ClassificationClassifier Stacking for Native Language Identification
This paper reports our contribution (team WLZ) to the NLI Shared Task 2017 (essay track). We first extract lexical and syntactic features from the essays, perform feature weighting and selection, and train linear support…
Language AcquisitionLanguage IdentificationNative Language IdentificationText Classification+1Native Language Identification using Phonetic Algorithms
In this paper, we discuss the results of the IUCL system in the NLI Shared Task 2017. For our system, we explore a variety of phonetic algorithms to generate features for Native Language Identification. These features ar…
Language IdentificationNative Language IdentificationA deep-learning based native-language classification by using a latent semantic analysis for the NLI Shared Task 2017
This paper proposes a deep-learning based native-language identification (NLI) using a latent semantic analysis (LSA) as a participant (ETRI-SLP) of the NLI Shared Task 2017 where the NLI Shared Task 2017 aims to detect …
Automatic Speech Recognition (ASR)Dimensionality ReductionGeneral ClassificationLanguage Identification+2Fusion of Simple Models for Native Language Identification
In this paper we describe the approaches we explored for the 2017 Native Language Identification shared task. We focused on simple word and sub-word units avoiding heavy use of hand-crafted features. Following recent tre…
Information RetrievalLanguage IdentificationNative Language IdentificationStacked Sentence-Document Classifier Approach for Improving Native Language Identification
In this paper, we describe the approach of the ItaliaNLP Lab team to native language identification and discuss the results we submitted as participants to the essay track of NLI Shared Task 2017. We introduce for the fi…
Document ClassificationLanguage IdentificationNative Language IdentificationSentence+2Vector Space Model as Cognitive Space for Text Classification
In this era of digitization, knowing the user's sociolect aspects have become essential features to build the user specific recommendation systems. These sociolect aspects could be found by mining the user's language sha…
Author ProfilingClassificationGender PredictionGeneral Classification+5Can string kernels pass the test of time in Native Language Identification?
We describe a machine learning approach for the 2017 shared task on Native Language Identification (NLI). The proposed approach combines several kernels using multiple kernel learning. While most of our kernels are based…
Language IdentificationNative Language IdentificationNative Language Identification on Text and Speech
This paper presents an ensemble system combining the output of multiple SVM classifiers to native language identification (NLI). The system was submitted to the NLI Shared Task 2017 fusion track which featured students e…
Language IdentificationNative Language IdentificationImproving Native Language Identification by Using Spelling Errors
In this paper, we explore spelling errors as a source of information for detecting the native language of a writer, a previously under-explored area. We note that character n-grams from misspelled words are very indicati…
Language IdentificationNative Language IdentificationLearning with learner corpora: Using the TLE for native language identification
Native Language Identification using Stacked Generalization
Ensemble methods using multiple classifiers have proven to be the most successful approach for the task of Native Language Identification (NLI), achieving the current state of the art. However, a systematic examination o…
Language IdentificationNative Language IdentificationAdvancing Linguistic Features and Insights by Label-informed Feature Grouping: An Exploration in the Context of Native Language Identification
We propose a hierarchical clustering approach designed to group linguistic features for supervised machine learning that is inspired by variationist linguistics. The method makes it possible to abstract away from the ind…
ClusteringLanguage AcquisitionLanguage IdentificationNative Language Identification+1