paper-with-me

홈 › Papers

Measure and Improve Robustness in NLP Models: A Survey

2021-12-15 · NAACL 2022 7 · Xuezhi Wang, Haohan Wang, Diyi Yang

As NLP models achieved state-of-the-art performances over benchmarks and gained wide applications, it has been increasingly important to ensure the safe deployment of these models in the real world, e.g., making sure the models are robust against unseen or challenging scenarios. Despite robustness being an increasingly studied topic, it has been separately explored in applications like vision and NLP, with various definitions, evaluation and mitigation strategies in multiple lines of research. In this paper, we aim to provide a unifying survey of how to define, measure and improve robustness in NLP. We first connect multiple definitions of robustness, then unify various lines of work on identifying robustness failures and evaluating models' robustness. Correspondingly, we present mitigation strategies that are data-driven, model-driven, and inductive-prior-based, with a more systematic view of how to effectively improve robustness in NLP models. Finally, we conclude by outlining open challenges and future directions to motivate further research in this area.

📄 PDF Abstract BibTeX arXiv:2112.08313

Code (0)

등록된 구현이 없습니다.

Tasks

Survey

Similar Papers 제목 키워드 기반

Are Large Language Models Chameleons? An Attempt to Simulate Social Surveys

2024-05-29 · Mingmeng Geng, Sihong He, Roberto Trotta

Can large language models (LLMs) simulate social surveys? To answer this question, we conducted millions of simulations in which LLMs were asked to answer subjective questions. A comparison of different LLM responses wit…

Survey

From Structural Equation Modelling to Double Machine Learning: Robustness Analysis for Survey-Based Research

2026-07-01 · Ka Ching Chan, Qiana Liu, Sanjib Tiwari, Ranga Chimhundu arxiv

Structural equation modelling (SEM) is widely used in survey-based business and information systems research to assess latent constructs and theory-driven structural relationships. However, SEM path significance is obtai…

EvalSafetyGap: A Hybrid Survey and Conceptual Framework for LLM Evaluation-Safety Failures

2026-06-29 · Buğra Alperen Uluırmak, Rifat Kurban arxiv

LLM evaluation and AI safety face a shared measurement problem: benchmark scores, reward-model signals, and reported safety metrics can improve while the latent properties they are meant to represent remain difficult to …

Adversarial Robustness

Prompt Robustness Is Task-Dependent: Comparing Objective and Belief-Style Questions in LLM Evaluation

2026-07-06 · Sadia Kamal, Arefa Patwary, Anthony Marchiafava, Atriya Sen 외 arxiv

Survey-style evaluations of large language models often treat a prompted response as a measure of a model's values or beliefs. This assumption is particularly fragile when responses are read as evidence of political valu…

A Survey of Neural Network Robustness Assessment in Image Recognition

2024-04-12 · Jie Wang, Jun Ai, Minyan Lu, Haoran Su 외

In recent years, there has been significant attention given to the robustness assessment of neural networks. Robustness plays a critical role in ensuring reliable operation of artificial intelligence (AI) systems in comp…

Adversarial Robustnessimage-classificationImage ClassificationSurvey