paper-with-me

Papers

RUPBench: Benchmarking Reasoning Under Perturbations for Robustness Evaluation in Large Language Models

2024-06-16 · Yuqing Wang, Yun Zhao

With the increasing use of large language models (LLMs), ensuring reliable performance in diverse, real-world environments is essential. Despite their remarkable achievements, LLMs often struggle with adversarial inputs, significantly impacting their effectiveness in practical applications. To systematically understand the robustness of LLMs, we present RUPBench, a comprehensive benchmark designed to evaluate LLM robustness across diverse reasoning tasks. Our benchmark incorporates 15 reasoning datasets, categorized into commonsense, arithmetic, logical, and knowledge-intensive reasoning, and introduces nine types of textual perturbations at lexical, syntactic, and semantic levels. By examining the performance of state-of-the-art LLMs such as GPT-4o, Llama3, Phi-3, and Gemma on both original and perturbed datasets, we provide a detailed analysis of their robustness and error patterns. Our findings highlight that larger models tend to exhibit greater robustness to perturbations. Additionally, common error types are identified through manual inspection, revealing specific challenges faced by LLMs in different reasoning contexts. This work provides insights into areas where LLMs need further improvement to handle diverse and noisy inputs effectively.

📄 PDF Abstract BibTeX arXiv:2406.11020

Code (1)

eternityyw/rupbench 공식 구현 pytorch

Tasks

Benchmarking

Similar Papers 제목 키워드 기반

Benchmarking Robustness of Multimodal Image-Text Models under Distribution Shift

2022-12-15 · JieLin Qiu, Yi Zhu, Xingjian Shi, Florian Wenzel 외

Multimodal image-text models have shown remarkable performance in the past few years. However, evaluating robustness against distribution shifts is crucial before adopting them in real-world applications. In this work, w…

BenchmarkingImage CaptioningImage GenerationImage-text Retrieval+6

Benchmarking Robustness of Deep Learning Classifiers Using Two-Factor Perturbation

2021-03-02 · Wei Dai, Daniel Berleant

This paper adds to the fundamental body of work on benchmarking the robustness of deep learning (DL) classifiers. We innovate a new benchmarking methodology to evaluate robustness of DL classifiers. Also, we introduce a …

BenchmarkingDeep LearningVocal Bursts Valence Prediction

An Investigation of Robustness of LLMs in Mathematical Reasoning: Benchmarking with Mathematically-Equivalent Transformation of Advanced Mathematical Problems

2025-08-12 · Yuren Hao, Xiang Wan, ChengXiang Zhai arxiv

In this paper, we introduce a systematic framework beyond conventional method to assess LLMs' mathematical-reasoning robustness by stress-testing them on advanced math problems that are mathematically equivalent but with…

Mathematical Reasoning

Benchmarking MLLM-based Web Understanding: Reasoning, Robustness and Safety

2025-09-26 · Junliang Liu, Jingyu Xiao, Wenxin Tang, Zhixian Wang 외 arxiv

Multimodal large language models (MLLMs) are increasingly deployed as the core reasoning engine for web-facing systems, powering GUI agents and front-end automation that must interpret page structure, select actionable w…

Code Generation

RL-Based Method for Benchmarking the Adversarial Resilience and Robustness of Deep Reinforcement Learning Policies

2019-06-03 · Vahid Behzadan, William Hsu

This paper investigates the resilience and robustness of Deep Reinforcement Learning (DRL) policies to adversarial perturbations in the state space. We first present an approach for the disentanglement of vulnerabilities…

BenchmarkingDeep Reinforcement LearningDisentanglementreinforcement-learning+4