paper-with-me

Papers

HellaSwag-Pro: A Large-Scale Bilingual Benchmark for Evaluating the Robustness of LLMs in Commonsense Reasoning

2025-02-17 · Xiaoyuan Li, Moxin Li, Rui Men, Yichang Zhang, Keqin Bao, Wenjie Wang, Fuli Feng, Dayiheng Liu, Junyang Lin

Large language models (LLMs) have shown remarkable capabilities in commonsense reasoning; however, some variations in questions can trigger incorrect responses. Do these models truly understand commonsense knowledge, or just memorize expression patterns? To investigate this question, we present the first extensive robustness evaluation of LLMs in commonsense reasoning. We introduce HellaSwag-Pro, a large-scale bilingual benchmark consisting of 11,200 cases, by designing and compiling seven types of question variants. To construct this benchmark, we propose a two-stage method to develop Chinese HellaSwag, a finely annotated dataset comprising 12,000 instances across 56 categories. We conduct extensive experiments on 41 representative LLMs, revealing that these LLMs are far from robust in commonsense reasoning. Furthermore, this robustness varies depending on the language in which the LLM is tested. This work establishes a high-quality evaluation benchmark, with extensive experiments offering valuable insights to the community in commonsense reasoning for LLMs.

📄 PDF Abstract BibTeX arXiv:2502.11393

Code (0)

등록된 구현이 없습니다.

Tasks

HellaSwag

Similar Papers 제목 키워드 기반

What the HellaSwag? On the Validity of Common-Sense Reasoning Benchmarks

2025-04-10 · Pavel Chizhov, Mattia Nee, Pierre-Carl Langlais, Ivan P. Yamshchikov

Common-sense reasoning is a key language model capability because it encapsulates not just specific factual knowledge but rather general language and world understanding. Measuring common-sense reasoning, therefore, is c…

Common Sense ReasoningHellaSwagModel Selection

SAGE: Scalable Automated Robustness Augmentation for LLM Knowledge Evaluation

2026-05-12 · Xiaoyuan Li, Yuzhe Wang, Moxin Li, Keqin Bao 외 arxiv

Large Language Models (LLMs) achieve strong performance on standard knowledge evaluation benchmarks, yet recent work shows that their knowledge capabilities remain brittle under question variants that test the same knowl…

Reinforcement Learning

BiToD: A Bilingual Multi-Domain Dataset For Task-Oriented Dialogue Modeling

2021-06-05 · Zhaojiang Lin, Andrea Madotto, Genta Indra Winata, Peng Xu 외

Task-oriented dialogue (ToD) benchmarks provide an important avenue to measure progress and develop better conversational agents. However, existing datasets for end-to-end ToD modeling are limited to a single language, h…

Cross-Lingual TransferTransfer Learning

Punching Above Their Weight: Classification-Head Fine-Tuning of Tiny Language Models (TLMs) for Verifiable Multiple-Choice Tasks

2026-07-04 · Bhavesh Sood, Jaromir Savelka arxiv

We define Tiny Language Models (TLMs) as models below roughly 3B parameters that fit on mainstream consumer devices. We study how to adapt them for and use them on verifiable multiple-choice tasks. We compare three LoRA-…

Towards Multilingual LLM Evaluation for European Languages

2024-10-11 · Klaudia Thellmann, Bernhard Stadler, Michael Fromm, Jasper Schulze Buschhoff 외

The rise of Large Language Models (LLMs) has revolutionized natural language processing across numerous languages and tasks. However, evaluating LLM performance in a consistent and meaningful way across multiple European…

ARCGSM8KHellaSwagMMLU+1