paper-with-me

Papers

ReXTime: A Benchmark Suite for Reasoning-Across-Time in Videos

2024-06-27 · Jr-Jen Chen, Yu-Chien Liao, Hsi-Che Lin, Yu-Chu Yu, Yen-Chun Chen, Yu-Chiang Frank Wang

We introduce ReXTime, a benchmark designed to rigorously test AI models' ability to perform temporal reasoning within video events. Specifically, ReXTime focuses on reasoning across time, i.e. human-like understanding when the question and its corresponding answer occur in different video segments. This form of reasoning, requiring advanced understanding of cause-and-effect relationships across video segments, poses significant challenges to even the frontier multimodal large language models. To facilitate this evaluation, we develop an automated pipeline for generating temporal reasoning question-answer pairs, significantly reducing the need for labor-intensive manual annotations. Our benchmark includes 921 carefully vetted validation samples and 2,143 test samples, each manually curated for accuracy and relevance. Evaluation results show that while frontier large language models outperform academic models, they still lag behind human performance by a significant 14.3% accuracy gap. Additionally, our pipeline creates a training dataset of 9,695 machine generated samples without manual effort, which empirical studies suggest can enhance the across-time reasoning via fine-tuning.

📄 PDF Abstract BibTeX arXiv:2406.19392

Code (1)

rextime/rextime 공식 구현

Similar Papers 제목 키워드 기반

EventCoT: Event-centric Video Chain-of-thought for Reasoning Temporal Localization

2026-07-06 · Youngkil Song, Yoonjae Baek, Dongwon Kim, Inho Kim 외 arxiv

Reasoning temporal localization (RTL) requires a model to generate an answer that itself contains the time interval supporting it, so high-level reasoning and precise temporal grounding must be produced jointly in a sing…

Video Question Answering

GraphThinker: Reinforcing Temporally Grounded Video Reasoning with Event Graph Thinking

2026-02-19 · Zixu Cheng, Da Li, Jian Hu, Yuhang Zang 외 arxiv

Video reasoning requires a fine-grained understanding of the temporal dependencies and event-level relations between objects and events in videos. Current Multimodal Large Language Models (MLLMs) are prone to severe temp…

Visual Grounding

Medmarks: A Comprehensive Open-Source LLM Benchmark Suite for Medical Tasks

2026-05-02 · Benjamin Warner, Ratna Sagari Grandhi, Max Kieffer, Aymane Ouraq 외 arxiv

Evaluating large language models (LLMs) for medical applications remains challenging due to benchmark saturation, limited data accessibility, and insufficient coverage of relevant tasks. Existing suites have either satur…

Reinforcement LearningInformation ExtractionQuestion Answering

GRE Suite: Geo-localization Inference via Fine-Tuned Vision-Language Models and Enhanced Reasoning Chains

2025-05-24 · Chun Wang, Xiaoran Pan, Zihao Pan, Haofan Wang 외

Recent advances in Visual Language Models (VLMs) have demonstrated exceptional performance in visual reasoning tasks. However, geo-localization presents unique challenges, requiring the extraction of multigranular visual…

geo-localizationVisual ReasoningWorld Knowledge

OpenExempt: A Diagnostic Benchmark for Legal Reasoning and a Framework for Creating Custom Benchmarks on Demand

2026-01-19 · Sergio Servantez, Sarah B. Lawsky, Rajiv Jain, Daniel W. Linna 외 arxiv

Reasoning benchmarks have played a crucial role in the progress of language models. Yet rigorous evaluation remains a significant challenge as static question-answer pairs provide only a snapshot of performance, compress…

Legal Reasoning