paper-with-me

Papers Linguistic Acceptability

“Linguistic Acceptability” 태그가 달린 논문 82편 · 필터 해제

Data filtering methods for training language models

2026-05-28 · Egor Shevchenko, Elena Bruches arxiv

Data quality is a critical factor in the effectiveness of machine learning models. Label errors, present even in widely used benchmarks, introduce noise into training data and reduce model generalization. In this work, w…

Linguistic AcceptabilityEmotion ClassificationLabel Error DetectionText Classification

DunbaaBERT: From Sacrifice to Semantics

2026-05-26 · Iffat Maab, Waleed Jamil, Raphael Schmitt arxiv

Large language models have achieved strong performance across many NLP tasks, yet Urdu remains comparatively underexplored due to limited resources and fragmented evaluation settings. To address this gap, we introduce Du…

Linguistic AcceptabilityNews ClassificationSentiment Analysis

Dialects of Translationese Shape Language Model Learning

2026-02-18 · Jenny Kunz arxiv

Machine-translated data is widely used in multilingual NLP, particularly where native text is scarce. However, translated text differs systematically from native text. This phenomenon is known as translationese, and it r…

Linguistic AcceptabilityLanguage Modelling

Preferences for Idiomatic Language are Acquired Slowly -- and Forgotten Quickly: A Case Study on Swedish

2026-02-03 · Jenny Kunz arxiv

In this study, we investigate how language models develop preferences for \textit{idiomatic} as compared to \textit{linguistically acceptable} Swedish, both during pretraining and when adapting a model from English to Sw…

Linguistic Acceptability

DaLA: Danish Linguistic Acceptability Evaluation Guided by Real World Errors

2025-12-04 · Gianluca Barmina, Nathalie Carmen Hau Norman, Peter Schneider-Kamp, Lukas Galke Poech arxiv

We present an enhanced benchmark for evaluating linguistic acceptability in Danish. We first analyze the most common errors found in written Danish. Based on this analysis, we introduce a set of fourteen corruption funct…

Linguistic Acceptability

SindBERT, the Sailor: Charting the Seas of Turkish NLP

2025-10-24 · Raphael Schmitt, Stefan Schweter arxiv

Transformer models have revolutionized NLP, yet many morphologically rich languages remain underrepresented in large-scale pre-training efforts. With SindBERT, we set out to chart the seas of Turkish NLP, providing the f…

Linguistic AcceptabilityPart-Of-Speech Tagging

\textsc{CantoNLU}: A benchmark for Cantonese natural language understanding

2025-10-23 · Junghyun Min, York Hay Ng, Sophia Chan, Helena Shunhua Zhao 외 arxiv

Cantonese, although spoken by millions, remains under-resourced due to policy and diglossia. To address this scarcity of evaluation frameworks for Cantonese, we introduce \textsc{\textbf{CantoNLU}}, a benchmark for Canto…

Natural Language UnderstandingNatural Language InferenceWord Sense DisambiguationLinguistic Acceptability

Family Matters: Language Transfer and Merging for Adapting Small LLMs to Faroese

2025-10-01 · Jenny Kunz, Iben Nyholm Debess, Annika Simonsen arxiv

We investigate strategies for adapting small, efficient language models to Faroese, a low-resource North Germanic language. Starting from English-pretrained models, we apply continued pre-training on related Scandinavian…

Linguistic AcceptabilityReading Comprehension

Tree Reward-Aligned Search for TReASURe in Masked Diffusion Language Models

2025-09-27 · Zichao Yu, Ming Li, Wenyi Zhang, Weiguo Gao arxiv

Tree search has recently emerged as a powerful framework for aligning generative models with task-specific rewards at test time. Applying tree search to Masked Diffusion Language Models, however, introduces two key chall…

Linguistic Acceptability

QFrCoLA: a Quebec-French Corpus of Linguistic Acceptability Judgments

2025-08-23 · David Beauchemin, Richard Khoury arxiv

Large and Transformer-based language models perform outstandingly in various downstream tasks. However, there is limited understanding regarding how these models internalize linguistic knowledge, so various linguistic be…

Linguistic AcceptabilityBinary Classification

Dissecting Bias in LLMs: A Mechanistic Interpretability Perspective

2025-06-05 · Bhavik Chandna, Zubair Bashir, Procheta Sen

Large Language Models (LLMs) are known to exhibit social, demographic, and gender biases, often as a consequence of the data on which they are trained. In this work, we adopt a mechanistic interpretability approach to an…

Linguistic Acceptabilitynamed-entity-recognitionNamed Entity Recognition

Fietje: An open, efficient LLM for Dutch

2024-12-19 · Bram Vanroy

This paper introduces Fietje, a family of small language models (SLMs) specifically designed for the Dutch language. The model is based on Phi 2, an English-centric model of 2.7 billion parameters. Fietje demonstrated co…

Linguistic AcceptabilitySentiment AnalysisWord Sense DisambiguationWorld Knowledge

Robust ASR Error Correction with Conservative Data Filtering

2024-07-18 · Takuma Udagawa, Masayuki Suzuki, Masayasu Muraoka, Gakuto Kurata

Error correction (EC) based on large language models is an emerging technology to enhance the performance of automatic speech recognition (ASR) systems. Generally, training data for EC are collected by automatically pair…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Linguistic Acceptabilityspeech-recognition+1

Learning Phonotactics from Linguistic Informants

2024-05-08 · Canaan Breiss, Alexis Ross, Amani Maina-Kilaas, Roger Levy 외

We propose an interactive approach to language learning that utilizes linguistic acceptability judgments from an informant (a competent language user) to learn a grammar. Given a grammar formalism and a framework for syn…

Linguistic Acceptability

MELA: Multilingual Evaluation of Linguistic Acceptability

2023-11-15 · Ziyin Zhang, Yikang Liu, Weifang Huang, Junyu Mao 외

In this work, we present the largest benchmark to date on linguistic acceptability: Multilingual Evaluation of Linguistic Acceptability -- MELA, with 46K samples covering 10 languages from a diverse set of language famil…

Code GenerationCross-Lingual TransferLinguistic AcceptabilityMulti-Task Learning+1

Data-Free Distillation of Language Model by Text-to-Text Transfer

2023-11-03 · Zheyuan Bai, Xinduo Liu, Hailin Hu, Tianyu Guo 외

Data-Free Knowledge Distillation (DFKD) plays a vital role in compressing the model when original training data is unavailable. Previous works for DFKD in NLP mainly focus on distilling encoder-only structures like BERT …

Data-free Knowledge DistillationDiversityKnowledge DistillationLanguage Modeling+5

Not all layers are equally as important: Every Layer Counts BERT

2023-11-03 · Lucas Georges Gabriel Charpentier, David Samuel

This paper introduces a novel modification of the transformer architecture, tailored for the data-efficient pretraining of language models. This aspect is evaluated by participating in the BabyLM challenge, where our sol…

AllLinguistic AcceptabilityNatural Language Inference

How well can machine-generated texts be identified and can language models be trained to avoid identification?

2023-10-25 · Sinclair Schneider, Florian Steuber, Joao A. G. Schneider, Gabi Dreo Rodosek

With the rise of generative pre-trained transformer models such as GPT-3, GPT-NeoX, or OPT, distinguishing human-generated texts from machine-generated ones has become important. We refined five separate language models …

Linguistic AcceptabilityText Generation

JCoLA: Japanese Corpus of Linguistic Acceptability

2023-09-22 · Taiga Someya, Yushi Sugimoto, Yohei Oseki

Neural language models have exhibited outstanding performance in a range of downstream tasks. However, there is limited understanding regarding the extent to which these models internalize syntactic knowledge, so that va…

ArticlesLinguistic Acceptability

Defense of Adversarial Ranking Attack in Text Retrieval: Benchmark and Baseline via Detection

2023-07-31 · Xuanang Chen, Ben He, Le Sun, Yingfei Sun

Neural ranking models (NRMs) have undergone significant development and have become integral components of information retrieval (IR) systems. Unfortunately, recent research has unveiled the vulnerability of NRMs to adve…

Adversarial AttackInformation RetrievalLinguistic AcceptabilityRetrieval+1
1–20 / 82 다음 →