paper-with-me

Papers

InterChart: Benchmarking Visual Reasoning Across Decomposed and Distributed Chart Information

2025-08-11 · Anirudh Iyengar Kaniyar Narayana Iyengar, Srija Mukhopadhyay, Adnan Qidwai, Shubhankar Singh, Dan Roth, Vivek Gupta arxiv

We introduce InterChart, a diagnostic benchmark that evaluates how well vision-language models (VLMs) reason across multiple related charts, a task central to real-world applications such as scientific reporting, financial analysis, and public policy dashboards. Unlike prior benchmarks focusing on isolated, visually uniform charts, InterChart challenges models with diverse question types ranging from entity inference and trend correlation to numerical estimation and abstract multi-step reasoning grounded in 2-3 thematically or structurally related charts. We organize the benchmark into three tiers of increasing difficulty: (1) factual reasoning over individual charts, (2) integrative analysis across synthetically aligned chart sets, and (3) semantic inference over visually complex, real-world chart pairs. Our evaluation of state-of-the-art open- and closed-source VLMs reveals consistent and steep accuracy declines as chart complexity increases. We find that models perform better when we decompose multi-entity charts into simpler visual units, underscoring their struggles with cross-chart integration. By exposing these systematic limitations, InterChart provides a rigorous framework for advancing multimodal reasoning in complex, multi-visual environments.

📄 PDF Abstract BibTeX arXiv:2508.07630

Code (0)

등록된 구현이 없습니다.

Tasks

Multimodal ReasoningVisual Reasoning

Similar Papers 제목 키워드 기반

DecompSR: A dataset for decomposed analyses of compositional multihop spatial reasoning

2025-11-04 · Lachlan McPheat, Navdeep Kaur, Robert Blackwell, Alessandra Russo 외 arxiv

We introduce DecompSR, decomposed spatial reasoning, a large benchmark dataset (over 5m datapoints) and generation framework designed to analyse compositional spatial reasoning ability. The generation of DecompSR allows …

Spatial Reasoning

PDCR: Perception-Decomposed Confidence Reward for Vision-Language Reasoning

2026-05-13 · Hee Suk Yoon, Eunseop Yoon, Ji Woo Hong, SooHwan Eom 외 arxiv

Reinforcement Learning with Verifiable Rewards (RLVR) traditionally relies on a sparse, outcome-based signal. Recent work shows that providing a fine-grained, model-intrinsic signal (rewarding the confidence growth in th…

Reinforcement Learning

UniTransfer: Video Concept Transfer via Progressive Spatial and Timestep Decomposition

2025-09-25 · Guojun Lei, Rong Zhang, Chi Wang, Tianhang Liu 외 arxiv

We propose a novel architecture UniTransfer, which introduces both spatial and diffusion timestep decomposition in a progressive paradigm, achieving precise and controllable video concept transfer. Specifically, in terms…

Representation Learning

Emergent Strategic Reasoning Risks in AI: A Taxonomy-Driven Evaluation Framework

2026-04-23 · Tharindu Kumarage, Lisa Bauer, Yao Ma, Dan Rosen 외 arxiv

As reasoning capacity and deployment scope grow in tandem, large language models (LLMs) gain the capacity to engage in behaviors that serve their own objectives, a class of risks we term Emergent Strategic Reasoning Risk…

Video Models Can Reason with Verifiable Rewards

2026-05-14 · Tinghui Zhu, Sheng Zhang, James Y. Huang, Selena Song 외 arxiv

Video diffusion models have made rapid progress in perceptual realism and temporal coherence, but they remain primarily optimized for plausible generation rather than verifiable reasoning. This limitation is especially p…

Reinforcement LearningVisual ReasoningVideo Generation