paper-with-me

Papers

Native Language Identification Using Large, Longitudinal Data

2014-05-01 · LREC 2014 5 · Xiao Jiang, Yufan Guo, Jeroen Geertzen, Dora Alexopoulou, Lin Sun, Anna Korhonen

Native Language Identification (NLI) is a task aimed at determining the native language (L1) of learners of second language (L2) on the basis of their written texts. To date, research on NLI has focused on relatively small corpora. We apply NLI to the recently released EFCamDat corpus which is not only multiple times larger than previous L2 corpora but also provides longitudinal data at several proficiency levels. Our investigation using accurate machine learning with a wide range of linguistic features reveals interesting patterns in the longitudinal data which are useful for both further development of NLI and its application to research on L2 acquisition.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

BIG-bench Machine LearningLanguage IdentificationNative Language IdentificationText Classification

Similar Papers 제목 키워드 기반

Adversarial Deep Learning in EEG Biometrics

2019-03-27 · Ozan Ozdenizci, Ye Wang, Toshiaki Koike-Akino, Deniz Erdogmus

Deep learning methods for person identification based on electroencephalographic (EEG) brain activity encounters the problem of exploiting the temporally correlated structures or recording session specific variability wi…

Deep LearningEEGElectroencephalogram (EEG)Person Identification+1

A Review on Generative AI Models for Synthetic Medical Text, Time Series, and Longitudinal Data

2024-11-19 · Mohammad Loni, Fatemeh Poursalim, Mehdi Asadi, Arash Gharehbaghi

This paper presents the results of a novel scoping review on the practical models for generating three different types of synthetic health records (SHRs): medical text, time series, and longitudinal data. The innovative …

ImputationTime Series

Deep Learning Approach for Clinical Risk Identification Using Transformer Modeling of Heterogeneous EHR Data

2025-11-06 · Anzhuo Xie, Wei-Chen Chang arxiv

This study proposes a Transformer-based longitudinal modeling method to address challenges in clinical risk classification with heterogeneous Electronic Health Record (EHR) data, including irregular temporal patterns, la…

Improving Language Identification of Accented Speech

2022-03-31 · Kunnar Kukk, Tanel Alumäe

Language identification from speech is a common preprocessing step in many spoken language processing systems. In recent years, this field has seen fast progress, mostly due to the use of self-supervised models pretraine…

Language Identificationspeech-recognitionSpeech RecognitionSpoken language identification

Identification of morphological fingerprint in perinatal brains using quasi-conformal mapping and contrastive learning

2023-11-25 · Boyang Wang, Weihao Zheng, Ying Wang, Zhe Zhang 외

The morphological fingerprint in the brain is capable of identifying the uniqueness of an individual. However, whether such individual patterns are present in perinatal brains, and which morphological attributes or corti…

Contrastive LearningData Augmentation