paper-with-me

홈 › Papers

Statistically Profiling Biases in Natural Language Reasoning Datasets and Models

2021-02-09 · Shanshan Huang, Kenny Q. Zhu

Recent work has indicated that many natural language understanding and reasoning datasets contain statistical cues that may be taken advantaged of by NLP models whose capability may thus be grossly overestimated. To discover the potential weakness in the models, some human-designed stress tests have been proposed but they are expensive to create and do not generalize to arbitrary models. We propose a light-weight and general statistical profiling framework, ICQ (I-See-Cue), which automatically identifies possible biases in any multiple-choice NLU datasets without the need to create any additional test cases, and further evaluates through blackbox testing the extent to which models may exploit these biases.

📄 PDF Abstract BibTeX arXiv:2102.04632

Code (0)

등록된 구현이 없습니다.

Tasks

Multiple-choiceNatural Language Understanding

Similar Papers 제목 키워드 기반

Mapping the Multilingual Margins: Intersectional Biases of Sentiment Analysis Systems in English, Spanish, and Arabic

2022-04-07 · LTEDI (ACL) 2022 5 · António Câmara, Nina Taneja, Tamjeed Azad, Emily Allaway 외

As natural language processing systems become more widespread, it is necessary to address fairness issues in their implementation and deployment to ensure that their negative impacts on society are understood and minimiz…

FairnessregressionSentiment Analysis

Examining the Influence of Political Bias on Large Language Model Performance in Stance Classification

2024-07-25 · Lynnette Hui Xian Ng, Iain Cruickshank, Roy Ka-Wei Lee

Large Language Models (LLMs) have demonstrated remarkable capabilities in executing tasks based on natural language queries. However, these models, trained on curated datasets, inherently embody biases ranging from racia…

ClassificationLanguage ModelingLanguage ModellingLarge Language Model+2

SPARK: Susceptibility-Guided Profiling and Steering of Latent Reasoning States in Large Language Models

2026-07-11 · Dongxu Zhang, Yiding Sun, Zihao Guo, Xiangyang Yang 외 arxiv

Reasoning failures in large language models (LLMs) are usually evaluated from final answers, but a wrong answer does not reveal why the model failed. The same incorrect output may reflect missing capability, an unstable …

Testing the effectiveness of saliency-based explainability in NLP using randomized survey-based experiments

2022-11-25 · Adel Rahimi, Shaurya Jain

As the applications of Natural Language Processing (NLP) in sensitive areas like Political Profiling, Review of Essays in Education, etc. proliferate, there is a great need for increasing transparency in NLP models to bu…

StaffPro: an LLM Agent for Joint Staffing and Profiling

2025-07-29 · Alessio Maritan arxiv

Large language model (LLM) agents integrate pre-trained LLMs with modular algorithmic components and have shown remarkable reasoning and decision-making abilities. In this work, we investigate their use for two tightly i…