Linguistic Acceptability
5개 벤치마크 · 논문 82편 · 이 태스크의 논문 보기 →
Benchmarks
Most implemented
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
RoBERTa: A Robustly Optimized BERT Pretraining Approach
Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer
ALBERT: A Lite BERT for Self-supervised Learning of Language Representations
DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter
Big Bird: Transformers for Longer Sequences
Papers
Data filtering methods for training language models
Data quality is a critical factor in the effectiveness of machine learning models. Label errors, present even in widely used benchmarks, introduce noise into training data and reduce model generalization. In this work, w…
Linguistic AcceptabilityEmotion ClassificationLabel Error DetectionText ClassificationDunbaaBERT: From Sacrifice to Semantics
Large language models have achieved strong performance across many NLP tasks, yet Urdu remains comparatively underexplored due to limited resources and fragmented evaluation settings. To address this gap, we introduce Du…
Linguistic AcceptabilityNews ClassificationSentiment AnalysisDialects of Translationese Shape Language Model Learning
Machine-translated data is widely used in multilingual NLP, particularly where native text is scarce. However, translated text differs systematically from native text. This phenomenon is known as translationese, and it r…
Linguistic AcceptabilityLanguage ModellingPreferences for Idiomatic Language are Acquired Slowly -- and Forgotten Quickly: A Case Study on Swedish
In this study, we investigate how language models develop preferences for \textit{idiomatic} as compared to \textit{linguistically acceptable} Swedish, both during pretraining and when adapting a model from English to Sw…
Linguistic AcceptabilityDaLA: Danish Linguistic Acceptability Evaluation Guided by Real World Errors
We present an enhanced benchmark for evaluating linguistic acceptability in Danish. We first analyze the most common errors found in written Danish. Based on this analysis, we introduce a set of fourteen corruption funct…
Linguistic AcceptabilitySindBERT, the Sailor: Charting the Seas of Turkish NLP
Transformer models have revolutionized NLP, yet many morphologically rich languages remain underrepresented in large-scale pre-training efforts. With SindBERT, we set out to chart the seas of Turkish NLP, providing the f…
Linguistic AcceptabilityPart-Of-Speech Tagging