paper-with-me

홈 › Papers

Dissecting vocabulary biases datasets through statistical testing and automated data augmentation for artifact mitigation in Natural Language Inference

2023-12-14 · Dat Thanh Nguyen

In recent years, the availability of large-scale annotated datasets, such as the Stanford Natural Language Inference and the Multi-Genre Natural Language Inference, coupled with the advent of pre-trained language models, has significantly contributed to the development of the natural language inference domain. However, these crowdsourced annotated datasets often contain biases or dataset artifacts, leading to overestimated model performance and poor generalization. In this work, we focus on investigating dataset artifacts and developing strategies to address these issues. Through the utilization of a novel statistical testing procedure, we discover a significant association between vocabulary distribution and text entailment classes, emphasizing vocabulary as a notable source of biases. To mitigate these issues, we propose several automatic data augmentation strategies spanning character to word levels. By fine-tuning the ELECTRA pre-trained language model, we compare the performance of boosted models with augmented data against their baseline counterparts. The experiments demonstrate that the proposed approaches effectively enhance model accuracy and reduce biases by up to 0.66% and 1.14%, respectively.

📄 PDF Abstract BibTeX arXiv:2312.08747

Code (1)

datngu/nli-artifacts 공식 구현

Tasks

Data AugmentationLanguage ModelingLanguage ModellingNatural Language Inference

Methods 이 논문이 사용한 방법론

Multi-Head Attention 설명 없음
Attention 설명 없음
Refunds@Expedia|||How do I get a full refund from Expedia? “How do I get a full refund from Expedia? How do I get a full refund from Expedia? – Call ☎️ +1-(888) 829 (0881) or +1-805-330-4056 or +1-805-330-4056 for Quick Help &…
Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
Dense Connections Dense Connections, or Fully Connected Connections, are a type of layer in a deep neural network that use a linear operation where every input is connected to every output…
Focus 설명 없음
Attention Dropout Attention Dropout is a type of dropout used in attention-based architectures, where elements are randomly dropped out of the…
Adam 설명 없음

Similar Papers 제목 키워드 기반

Over-representation of phonological features in basic vocabulary doesn't replicate when controlling for spatial and phylogenetic effects

2025-12-08 · Frederic Blum arxiv

The statistical over-representation of phonological features in the basic vocabulary of languages is often interpreted as reflecting potentially universal sound symbolic patterns. However, most of those results have not …

Statistically Profiling Biases in Natural Language Reasoning Datasets and Models

2021-02-09 · Shanshan Huang, Kenny Q. Zhu

Recent work has indicated that many natural language understanding and reasoning datasets contain statistical cues that may be taken advantaged of by NLP models whose capability may thus be grossly overestimated. To disc…

Multiple-choiceNatural Language Understanding

Dissecting Bias in LLMs: A Mechanistic Interpretability Perspective

2025-06-05 · Bhavik Chandna, Zubair Bashir, Procheta Sen

Large Language Models (LLMs) are known to exhibit social, demographic, and gender biases, often as a consequence of the data on which they are trained. In this work, we adopt a mechanistic interpretability approach to an…

Linguistic Acceptabilitynamed-entity-recognitionNamed Entity Recognition

Dissecting a Small Artificial Neural Network

2025-01-03 · Xiguang Yang, Krish Arora, Michael Bachmann

We investigate the loss landscape and backpropagation dynamics of convergence for the simplest possible artificial neural network representing the logical exclusive-OR (XOR) gate. Cross-sections of the loss landscape in …

Language Integration in Fine-Tuning Multimodal Large Language Models for Image-Based Regression

2025-07-20 · Roy H. Jennings, Genady Paikin, Roy Shaul, Evgeny Soloveichik

Multimodal Large Language Models (MLLMs) show promise for image-based regression tasks, but current approaches face key limitations. Recent methods fine-tune MLLMs using preset output vocabularies and generic task-level …

Aesthetics Quality AssessmentNo-Reference Image Quality Assessmentregression