Native Language Identification Using Large, Longitudinal Data
Native Language Identification (NLI) is a task aimed at determining the native language (L1) of learners of second language (L2) on the basis of their written texts. To date, research on NLI has focused on relatively small corpora. We apply NLI to the recently released EFCamDat corpus which is not only multiple times larger than previous L2 corpora but also provides longitudinal data at several proficiency levels. Our investigation using accurate machine learning with a wide range of linguistic features reveals interesting patterns in the longitudinal data which are useful for both further development of NLI and its application to research on L2 acquisition.
Code (0)
등록된 구현이 없습니다.
Tasks
BIG-bench Machine LearningLanguage IdentificationNative Language IdentificationText ClassificationSimilar Papers 제목 키워드 기반
Adversarial Deep Learning in EEG Biometrics
Deep learning methods for person identification based on electroencephalographic (EEG) brain activity encounters the problem of exploiting the temporally correlated structures or recording session specific variability wi…
Deep LearningEEGElectroencephalogram (EEG)Person Identification+1A Review on Generative AI Models for Synthetic Medical Text, Time Series, and Longitudinal Data
This paper presents the results of a novel scoping review on the practical models for generating three different types of synthetic health records (SHRs): medical text, time series, and longitudinal data. The innovative …
ImputationTime SeriesDeep Learning Approach for Clinical Risk Identification Using Transformer Modeling of Heterogeneous EHR Data
This study proposes a Transformer-based longitudinal modeling method to address challenges in clinical risk classification with heterogeneous Electronic Health Record (EHR) data, including irregular temporal patterns, la…
Improving Language Identification of Accented Speech
Language identification from speech is a common preprocessing step in many spoken language processing systems. In recent years, this field has seen fast progress, mostly due to the use of self-supervised models pretraine…
Language Identificationspeech-recognitionSpeech RecognitionSpoken language identificationIdentification of morphological fingerprint in perinatal brains using quasi-conformal mapping and contrastive learning
The morphological fingerprint in the brain is capable of identifying the uniqueness of an individual. However, whether such individual patterns are present in perinatal brains, and which morphological attributes or corti…
Contrastive LearningData Augmentation