paper-with-me

Papers

Benchmarking Critical Questions Generation: A Challenging Reasoning Task for Large Language Models

2025-05-16 · Banca Calvo Figueras, Rodrigo Agerri

The task of Critical Questions Generation (CQs-Gen) aims to foster critical thinking by enabling systems to generate questions that expose underlying assumptions and challenge the validity of argumentative reasoning structures. Despite growing interest in this area, progress has been hindered by the lack of suitable datasets and automatic evaluation standards. This paper presents a comprehensive approach to support the development and benchmarking of systems for this task. We construct the first large-scale dataset including $~$5K manually annotated questions. We also investigate automatic evaluation methods and propose a reference-based technique using large language models (LLMs) as the strategy that best correlates with human judgments. Our zero-shot evaluation of 11 LLMs establishes a strong baseline while showcasing the difficulty of the task. Data and code plus a public leaderboard are provided to encourage further research not only in terms of model performance, but also to explore the practical benefits of CQs-Gen for both automated reasoning and human critical thinking.

📄 PDF Abstract BibTeX arXiv:2505.11341

Code (0)

등록된 구현이 없습니다.

Tasks

Benchmarking

Similar Papers 제목 키워드 기반

CareMedEval dataset: Evaluating Critical Appraisal and Reasoning in the Biomedical Field

2025-11-05 · Doria Bonzi, Alexandre Guiggi, Frédéric Béchet, Carlos Ramisch 외 arxiv

Critical appraisal of scientific literature is an essential skill in the biomedical field. While large language models (LLMs) can offer promising support in this task, their reliability remains limited, particularly for …

R2I-Bench: Benchmarking Reasoning-Driven Text-to-Image Generation

2025-05-29 · Kaijie Chen, Zihao Lin, Zhiyang Xu, Ying Shen 외

Reasoning is a fundamental capability often required in real-world text-to-image (T2I) generation, e.g., generating ``a bitten apple that has been left in the air for more than a week`` necessitates understanding tempora…

BenchmarkingImage GenerationText to Image GenerationText-to-Image Generation

Seek-and-Solve: Benchmarking MLLMs for Visual Clue-Driven Reasoning in Daily Scenarios

2026-04-15 · Xiaomin Li, Tala Wang, Zichen Zhong, Ying Zhang 외 arxiv

Daily scenarios are characterized by visual richness, requiring Multimodal Large Language Models (MLLMs) to filter noise and identify decisive visual clues for accurate reasoning. Yet, current benchmarks predominantly ai…

SciVideoBench: Benchmarking Scientific Video Reasoning in Large Multimodal Models

2025-10-09 · Andong Deng, Taojiannan Yang, Shoubin Yu, Lincoln Spencer 외 arxiv

Large Multimodal Models (LMMs) have achieved remarkable progress across various capabilities; however, complex video reasoning in the scientific domain remains a significant and challenging frontier. Current video benchm…

Logical ReasoningVisual Grounding

V-REX: Benchmarking Exploratory Visual Reasoning via Chain-of-Questions

2025-12-12 · Chenrui Fan, Yijun Liang, Shweta Bhardwaj, Kwesi Cobbina 외 arxiv

While many vision-language models (VLMs) are developed to answer well-defined, straightforward questions with highly specified targets, as in most benchmarks, they often struggle in practice with complex open-ended tasks…

Visual Reasoning