paper-with-me

홈 › Papers

RESPONSE: Benchmarking the Ability of Language Models to Undertake Commonsense Reasoning in Crisis Situation

2025-03-14 · Aissatou Diallo, Antonis Bikakis, Luke Dickens, Anthony Hunter, Rob Miller

An interesting class of commonsense reasoning problems arises when people are faced with natural disasters. To investigate this topic, we present \textsf{RESPONSE}, a human-curated dataset containing 1789 annotated instances featuring 6037 sets of questions designed to assess LLMs' commonsense reasoning in disaster situations across different time frames. The dataset includes problem descriptions, missing resources, time-sensitive solutions, and their justifications, with a subset validated by environmental engineers. Through both automatic metrics and human evaluation, we compare LLM-generated recommendations against human responses. Our findings show that even state-of-the-art models like GPT-4 achieve only 37\% human-evaluated correctness for immediate response actions, highlighting significant room for improvement in LLMs' ability for commonsense reasoning in crises.

📄 PDF Abstract BibTeX arXiv:2503.11348

Code (0)

등록된 구현이 없습니다.

Tasks

Benchmarking

Methods 이 논문이 사용한 방법론

Attention 설명 없음
Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
Label Smoothing Label Smoothing is a regularization technique that introduces noise for the labels. This accounts for the fact that datasets may have mistakes in them, so maximizing the…
Position-Wise Feed-Forward Layer 설명 없음
Layer Normalization Unlike batch normalization, Layer Normalization directly estimates the normalization statistics from the summed inputs…
BPE Byte Pair Encoding, or BPE, is a subword segmentation algorithm that encodes rare and unknown words as sequences of subword units. The intuition is that various word…
Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
Multi-Head Attention 설명 없음

Similar Papers 제목 키워드 기반

Commonsense-Focused Dialogues for Response Generation: An Empirical Study

2021-09-14 · SIGDIAL (ACL) 2021 7 · Pei Zhou, Karthik Gopalakrishnan, Behnam Hedayatnia, Seokhwan Kim 외

Smooth and effective communication requires the ability to perform latent or explicit commonsense inference. Prior commonsense reasoning benchmarks (such as SocialIQA and CommonsenseQA) mainly focus on the discriminative…

Response GenerationText Generation

Gemini in Reasoning: Unveiling Commonsense in Multimodal Large Language Models

2023-12-29 · Yuqing Wang, Yun Zhao

The burgeoning interest in Multimodal Large Language Models (MLLMs), such as OpenAI's GPT-4V(ision), has significantly impacted both academic and industrial realms. These models enhance Large Language Models (LLMs) with …

HellaSwag

Shortcutted Commonsense: Data Spuriousness in Deep Learning of Commonsense Reasoning

2021-11-01 · EMNLP 2021 11 · Ruben Branco, António Branco, João António Rodrigues, João Ricardo Silva

Commonsense is a quintessential human capacity that has been a core challenge to Artificial Intelligence since its inception. Impressive results in Natural Language Processing tasks, including in commonsense reasoning, h…

Deep Learning

Sibyl: Empowering Empathetic Dialogue Generation in Large Language Models via Sensible and Visionary Commonsense Inference

2023-11-26 · Lanrui Wang, Jiangnan Li, Chenxu Yang, Zheng Lin 외

Recently, there has been a heightened interest in building chatbots based on Large Language Models (LLMs) to emulate human-like qualities in multi-turn conversations. Despite having access to commonsense knowledge to bet…

Dialogue Generation

SYNDICOM: Improving Conversational Commonsense with Error-Injection and Natural Language Feedback

2023-09-18 · Christopher Richardson, Anirudh Sundar, Larry Heck

Commonsense reasoning is a critical aspect of human communication. Despite recent advances in conversational AI driven by large language models, commonsense reasoning remains a challenging task. In this work, we introduc…

Response Generationvalid