paper-with-me

홈 › Papers

SWAG: A Large-Scale Adversarial Dataset for Grounded Commonsense Inference

2018-08-16 · EMNLP 2018 10 · Rowan Zellers, Yonatan Bisk, Roy Schwartz, Yejin Choi

Given a partial description like "she opened the hood of the car," humans can reason about the situation and anticipate what might come next ("then, she examined the engine"). In this paper, we introduce the task of grounded commonsense inference, unifying natural language inference and commonsense reasoning. We present SWAG, a new dataset with 113k multiple choice questions about a rich spectrum of grounded situations. To address the recurring challenges of the annotation artifacts and human biases found in many existing datasets, we propose Adversarial Filtering (AF), a novel procedure that constructs a de-biased dataset by iteratively training an ensemble of stylistic classifiers, and using them to filter the data. To account for the aggressive adversarial filtering, we use state-of-the-art language models to massively oversample a diverse set of potential counterfactuals. Empirical results demonstrate that while humans can solve the resulting inference problems with high accuracy (88%), various competitive models struggle on our task. We provide comprehensive analysis that indicates significant opportunities for future research.

📄 PDF Abstract BibTeX arXiv:1808.05326

Code (0)

등록된 구현이 없습니다.

Tasks

Common Sense ReasoningMultiple-choiceNatural Language InferenceQuestion Answering

Similar Papers 제목 키워드 기반

HellaSwag: Can a Machine Really Finish Your Sentence?

2019-05-19 · ACL 2019 7 · Rowan Zellers, Ari Holtzman, Yonatan Bisk, Ali Farhadi 외

Recent work by Zellers et al. (2018) introduced a new task of commonsense natural language inference: given an event description such as "A woman sits at a piano," a machine must select the most likely followup: "She set…

HellaSwagNatural Language InferenceSentenceSentence Completion

HellaSwag-Pro: A Large-Scale Bilingual Benchmark for Evaluating the Robustness of LLMs in Commonsense Reasoning

2025-02-17 · Xiaoyuan Li, Moxin Li, Rui Men, Yichang Zhang 외

Large language models (LLMs) have shown remarkable capabilities in commonsense reasoning; however, some variations in questions can trigger incorrect responses. Do these models truly understand commonsense knowledge, or …

HellaSwag

XDBERT: Distilling Visual Information to BERT from Cross-Modal Systems to Improve Language Understanding

2022-04-15 · ACL 2022 5 · Chan-Jan Hsu, Hung-Yi Lee, Yu Tsao

Transformer-based models are widely used in natural language understanding (NLU) tasks, and multimodal transformers have been effective in visual-language tasks. This study explores distilling visual information from pre…

Natural Language Understanding

PerCoR: Evaluating Commonsense Reasoning in Persian via Multiple-Choice Sentence Completion

2025-10-26 · Morteza Alikhani, Mohammadtaha Bagherifard, Erfan Zinvandi, Mehran Sarmadi arxiv

We introduced PerCoR (Persian Commonsense Reasoning), the first large-scale Persian benchmark for commonsense reasoning. PerCoR contains 106K multiple-choice sentence-completion problems drawn from more than forty news, …

Sentence Completion

Vokenization: Improving Language Understanding with Contextualized, Visual-Grounded Supervision

2020-10-14 · EMNLP 2020 11 · Hao Tan, Mohit Bansal

Humans learn language by listening, speaking, writing, reading, and also, via interaction with the multimodal real world. Existing language pre-training frameworks show the effectiveness of text-only self-supervision whi…

Image CaptioningLanguage ModelingLanguage Modelling