paper-with-me

홈 › Papers

EQUATE: A Benchmark Evaluation Framework for Quantitative Reasoning in Natural Language Inference

2019-01-11 · CONLL 2019 11 · Abhilasha Ravichander, Aakanksha Naik, Carolyn Rose, Eduard Hovy

Quantitative reasoning is a higher-order reasoning skill that any intelligent natural language understanding system can reasonably be expected to handle. We present EQUATE (Evaluating Quantitative Understanding Aptitude in Textual Entailment), a new framework for quantitative reasoning in textual entailment. We benchmark the performance of 9 published NLI models on EQUATE, and find that on average, state-of-the-art methods do not achieve an absolute improvement over a majority-class baseline, suggesting that they do not implicitly learn to reason with quantities. We establish a new baseline Q-REAS that manipulates quantities symbolically. In comparison to the best performing NLI model, it achieves success on numerical reasoning tests (+24.2%), but has limited verbal reasoning capabilities (-8.1%). We hope our evaluation framework will support the development of models of quantitative reasoning in language understanding.

📄 PDF Abstract BibTeX arXiv:1901.03735

Code (1)

AbhilashaRavichander/EQUATE 공식 구현

Tasks

Natural Language InferenceNatural Language Understanding

Similar Papers 제목 키워드 기반

MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding

2024-06-20 · Xinyu Fang, Kangrui Mao, Haodong Duan, Xiangyu Zhao 외

The advent of large vision-language models (LVLMs) has spurred research into their applications in multi-modal contexts, particularly in video understanding. Traditional VideoQA benchmarks, despite providing quantitative…

FormVideo Understanding

Design, Results and Industry Implications of the World's First Insurance Large Language Model Evaluation Benchmark

2025-11-11 · Hua Zhou, Bing Ma, Yufei Zhang, Yi Zhao arxiv

This paper comprehensively elaborates on the construction methodology, multi-dimensional evaluation system, and underlying design philosophy of CUFEInse v1.0. Adhering to the principles of "quantitative-oriented, expert-…

Domain Adaptation

Unified Embodied VLM Reasoning with Robotic Action via Autoregressive Discretized Pre-training

2025-12-30 · Yi Liu, Sukai Wang, Dafeng Wei, Xiaowei Cai 외 arxiv

General-purpose robotic systems operating in open-world environments must achieve both broad generalization and high-precision action execution, a combination that remains challenging for existing Vision-Language-Action …

Continuous Control

PsychBench: A comprehensive and professional benchmark for evaluating the performance of LLM-assisted psychiatric clinical practice

2025-02-28 · Shuyu Liu, Ruoxi Wang, Ling Zhang, Xuequan Zhu 외

The advent of Large Language Models (LLMs) offers potential solutions to address problems such as shortage of medical resources and low diagnostic consistency in psychiatric clinical practice. Despite this potential, a r…

BenchmarkingDiagnostic

Revisiting Reliability in the Reasoning-based Pose Estimation Benchmark

2025-07-17 · Junsu Kim, Naeun Kim, Jaeho Lee, Incheol Park 외

The reasoning-based pose estimation (RPE) benchmark has emerged as a widely adopted evaluation standard for pose-aware multimodal large language models (MLLMs). Despite its significance, we identified critical reproducib…

Multimodal ReasoningPose Estimation