Papers Linguistic Acceptability
“Linguistic Acceptability” 태그가 달린 논문 82편 · 필터 해제
Data filtering methods for training language models
Data quality is a critical factor in the effectiveness of machine learning models. Label errors, present even in widely used benchmarks, introduce noise into training data and reduce model generalization. In this work, w…
Linguistic AcceptabilityEmotion ClassificationLabel Error DetectionText ClassificationDunbaaBERT: From Sacrifice to Semantics
Large language models have achieved strong performance across many NLP tasks, yet Urdu remains comparatively underexplored due to limited resources and fragmented evaluation settings. To address this gap, we introduce Du…
Linguistic AcceptabilityNews ClassificationSentiment AnalysisDialects of Translationese Shape Language Model Learning
Machine-translated data is widely used in multilingual NLP, particularly where native text is scarce. However, translated text differs systematically from native text. This phenomenon is known as translationese, and it r…
Linguistic AcceptabilityLanguage ModellingPreferences for Idiomatic Language are Acquired Slowly -- and Forgotten Quickly: A Case Study on Swedish
In this study, we investigate how language models develop preferences for \textit{idiomatic} as compared to \textit{linguistically acceptable} Swedish, both during pretraining and when adapting a model from English to Sw…
Linguistic AcceptabilityDaLA: Danish Linguistic Acceptability Evaluation Guided by Real World Errors
We present an enhanced benchmark for evaluating linguistic acceptability in Danish. We first analyze the most common errors found in written Danish. Based on this analysis, we introduce a set of fourteen corruption funct…
Linguistic AcceptabilitySindBERT, the Sailor: Charting the Seas of Turkish NLP
Transformer models have revolutionized NLP, yet many morphologically rich languages remain underrepresented in large-scale pre-training efforts. With SindBERT, we set out to chart the seas of Turkish NLP, providing the f…
Linguistic AcceptabilityPart-Of-Speech Tagging\textsc{CantoNLU}: A benchmark for Cantonese natural language understanding
Cantonese, although spoken by millions, remains under-resourced due to policy and diglossia. To address this scarcity of evaluation frameworks for Cantonese, we introduce \textsc{\textbf{CantoNLU}}, a benchmark for Canto…
Natural Language UnderstandingNatural Language InferenceWord Sense DisambiguationLinguistic AcceptabilityFamily Matters: Language Transfer and Merging for Adapting Small LLMs to Faroese
We investigate strategies for adapting small, efficient language models to Faroese, a low-resource North Germanic language. Starting from English-pretrained models, we apply continued pre-training on related Scandinavian…
Linguistic AcceptabilityReading ComprehensionTree Reward-Aligned Search for TReASURe in Masked Diffusion Language Models
Tree search has recently emerged as a powerful framework for aligning generative models with task-specific rewards at test time. Applying tree search to Masked Diffusion Language Models, however, introduces two key chall…
Linguistic AcceptabilityQFrCoLA: a Quebec-French Corpus of Linguistic Acceptability Judgments
Large and Transformer-based language models perform outstandingly in various downstream tasks. However, there is limited understanding regarding how these models internalize linguistic knowledge, so various linguistic be…
Linguistic AcceptabilityBinary ClassificationDissecting Bias in LLMs: A Mechanistic Interpretability Perspective
Large Language Models (LLMs) are known to exhibit social, demographic, and gender biases, often as a consequence of the data on which they are trained. In this work, we adopt a mechanistic interpretability approach to an…
Linguistic Acceptabilitynamed-entity-recognitionNamed Entity RecognitionFietje: An open, efficient LLM for Dutch
This paper introduces Fietje, a family of small language models (SLMs) specifically designed for the Dutch language. The model is based on Phi 2, an English-centric model of 2.7 billion parameters. Fietje demonstrated co…
Linguistic AcceptabilitySentiment AnalysisWord Sense DisambiguationWorld KnowledgeRobust ASR Error Correction with Conservative Data Filtering
Error correction (EC) based on large language models is an emerging technology to enhance the performance of automatic speech recognition (ASR) systems. Generally, training data for EC are collected by automatically pair…
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Linguistic Acceptabilityspeech-recognition+1Learning Phonotactics from Linguistic Informants
We propose an interactive approach to language learning that utilizes linguistic acceptability judgments from an informant (a competent language user) to learn a grammar. Given a grammar formalism and a framework for syn…
Linguistic AcceptabilityMELA: Multilingual Evaluation of Linguistic Acceptability
In this work, we present the largest benchmark to date on linguistic acceptability: Multilingual Evaluation of Linguistic Acceptability -- MELA, with 46K samples covering 10 languages from a diverse set of language famil…
Code GenerationCross-Lingual TransferLinguistic AcceptabilityMulti-Task Learning+1Data-Free Distillation of Language Model by Text-to-Text Transfer
Data-Free Knowledge Distillation (DFKD) plays a vital role in compressing the model when original training data is unavailable. Previous works for DFKD in NLP mainly focus on distilling encoder-only structures like BERT …
Data-free Knowledge DistillationDiversityKnowledge DistillationLanguage Modeling+5Not all layers are equally as important: Every Layer Counts BERT
This paper introduces a novel modification of the transformer architecture, tailored for the data-efficient pretraining of language models. This aspect is evaluated by participating in the BabyLM challenge, where our sol…
AllLinguistic AcceptabilityNatural Language InferenceHow well can machine-generated texts be identified and can language models be trained to avoid identification?
With the rise of generative pre-trained transformer models such as GPT-3, GPT-NeoX, or OPT, distinguishing human-generated texts from machine-generated ones has become important. We refined five separate language models …
Linguistic AcceptabilityText GenerationJCoLA: Japanese Corpus of Linguistic Acceptability
Neural language models have exhibited outstanding performance in a range of downstream tasks. However, there is limited understanding regarding the extent to which these models internalize syntactic knowledge, so that va…
ArticlesLinguistic AcceptabilityDefense of Adversarial Ranking Attack in Text Retrieval: Benchmark and Baseline via Detection
Neural ranking models (NRMs) have undergone significant development and have become integral components of information retrieval (IR) systems. Unfortunately, recent research has unveiled the vulnerability of NRMs to adve…
Adversarial AttackInformation RetrievalLinguistic AcceptabilityRetrieval+1