Papers Native Language Identification
“Native Language Identification” 태그가 달린 논문 88편 · 필터 해제
The Impact of Editorial Intervention on Detecting Native Language Traces
Native Language Identification (NLI) is the task of determining an author's native language (L1) from their non-native writing. With the advent of human-AI co-authorship, learner texts are routinely corrected and rewritt…
Native Language IdentificationGrammatical Error CorrectionCan We Still Hear the Accent? Investigating the Resilience of Native Language Signals in the LLM Era
The evolution of writing assistance tools from machine translation to large language models (LLMs) has changed how researchers write. This study investigates whether this shift is homogenizing research papers by analyzin…
Native Language IdentificationMachine TranslationI can tell whether you are a Native Hawlêri Speaker! How ANN, CNN, and RNN perform in NLI-Native Language Identification
Native Language Identification (NLI) is a task in Natural Language Processing (NLP) that typically determines the native language of an author through their writing or a speaker through their speaking. It has various app…
Native Language IdentificationNLP Privacy Risk Identification in Social Media (NLP-PRISM): A Survey
Natural Language Processing (NLP) is integral to social media analytics but often processes content containing Personally Identifiable Information (PII), behavioral cues, and metadata raising privacy risks such as survei…
Native Language IdentificationSentiment AnalysisRobust Native Language Identification through Agentic Decomposition
Large language models (LLMs) often achieve high performance in native language identification (NLI) benchmarks by leveraging superficial contextual clues such as names, locations, and cultural stereotypes, rather than th…
Native Language IdentificationLeveraging Open-Source Large Language Models for Native Language Identification
Native Language Identification (NLI) - the task of identifying the native language (L1) of a person based on their writing in the second language (L2) - has applications in forensics, marketing, and second language acqui…
Feature EngineeringLanguage AcquisitionLanguage IdentificationMarketing+2Native Language Identification with Large Language Models
We present the first experiments on Native Language Identification (NLI) using LLMs such as GPT-4. NLI is the task of predicting a writer's first language by analyzing their writings in a second language, and is used in …
Language AcquisitionLanguage IdentificationNative Language IdentificationNative Language Identification with Big Bird Embeddings
Native Language Identification (NLI) intends to classify an author's native language based on their writing in another language. Historically, the task has heavily relied on time-consuming linguistic feature engineering,…
Computational EfficiencyFeature EngineeringLanguage IdentificationNative Language IdentificationTurkish Native Language Identification
In this paper, we present the first application of Native Language Identification (NLI) for the Turkish language. NLI involves predicting the writer's first language by analysing their writing in different languages. Whi…
Language IdentificationNative Language IdentificationScaling Native Language Identification with Transformer Adapters
Native language identification (NLI) is the task of automatically identifying the native language (L1) of an individual based on their language production in a learned language. It is useful for a variety of purposes inc…
Language IdentificationMarketingMulti-Label ClassificationMUlTI-LABEL-ClASSIFICATION+1Unravelling Interlanguage Facts via Explainable Machine Learning
Native language identification (NLI) is the task of training (via supervised machine learning) a classifier that guesses the native language of the author of a text. This task has been extensively researched in the last …
BIG-bench Machine LearningLanguage IdentificationNative Language IdentificationNative Language Identification and Reconstruction of Native Language Relationship Using Japanese Learner Corpus
A Deep Generative Approach to Native Language Identification
Native language identification (NLI) {--} identifying the native language (L1) of a person based on his/her writing in the second language (L2) {--} is useful for a variety of purposes, including marketing, security, and…
BIG-bench Machine LearningLanguage IdentificationLanguage ModellingMarketing+2Native-Language Identification with Attention
The paper explores how an attention-based approach can increase performance on the task of native-language identification (NLI), i.e., to identify an author’s first language given information expressed in a second langua…
Language IdentificationNative Language IdentificationInvestigating the effect of auxiliary objectives for the automated grading of learner English speech transcriptions
We address the task of automatically grading the language proficiency of spontaneous speech based on textual features from automatic speech recognition transcripts. Motivated by recent advances in multi-task learning, we…
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Language IdentificationLanguage Modeling+5A Report on the 2020 VUA and TOEFL Metaphor Detection Shared Task
In this paper, we report on the shared task on metaphor identification on VU Amsterdam Metaphor Corpus and on a subset of the TOEFL Native Language Identification Corpus. The shared task was conducted as apart of the ACL…
Language IdentificationNative Language IdentificationTopics to Avoid: Demoting Latent Confounds in Text Classification
Despite impressive performance on many text classification tasks, deep neural networks tend to learn frequent superficial patterns that are specific to the training data and do not always generalize well. In this work, w…
ClassificationGeneral ClassificationLanguage IdentificationNative Language Identification+2Towards Ethical Content-Based Detection of Online Influence Campaigns
The detection of clandestine efforts to influence users in online communities is a challenging problem with significant active development. We demonstrate that features derived from the text of user comments are useful f…
Language IdentificationNative Language IdentificationSentenceRegression or classification? Automated Essay Scoring for Norwegian
In this paper we present first results for the task of Automated Essay Scoring for Norwegian learner language. We analyze a number of properties of this task experimentally and assess (i) the formulation of the task as e…
Automated Essay ScoringBIG-bench Machine LearningClassificationGeneral Classification+4Anglicized Words and Misspelled Cognates in Native Language Identification
In this paper, we present experiments that estimate the impact of specific lexical choices of people writing in a second language (L2). In particular, we look at misspelled words that indicate lexical uncertainty on the …
Language IdentificationNative Language Identification