paper-with-me

홈 › Papers

RealHiTBench: A Comprehensive Realistic Hierarchical Table Benchmark for Evaluating LLM-Based Table Analysis

2025-06-16 · Pengzuo Wu, Yuhang Yang, Guangcheng Zhu, Chao Ye, Hong Gu, Xu Lu, Ruixuan Xiao, Bowen Bao, Yijing He, Liangyu Zha, Wentao Ye, Junbo Zhao, Haobo Wang

With the rapid advancement of Large Language Models (LLMs), there is an increasing need for challenging benchmarks to evaluate their capabilities in handling complex tabular data. However, existing benchmarks are either based on outdated data setups or focus solely on simple, flat table structures. In this paper, we introduce RealHiTBench, a comprehensive benchmark designed to evaluate the performance of both LLMs and Multimodal LLMs (MLLMs) across a variety of input formats for complex tabular data, including LaTeX, HTML, and PNG. RealHiTBench also includes a diverse collection of tables with intricate structures, spanning a wide range of task types. Our experimental results, using 25 state-of-the-art LLMs, demonstrate that RealHiTBench is indeed a challenging benchmark. Moreover, we also develop TreeThinker, a tree-based pipeline that organizes hierarchical headers into a tree structure for enhanced tabular reasoning, validating the importance of improving LLMs' perception of table hierarchies. We hope that our work will inspire further research on tabular data reasoning and the development of more robust models. The code and data are available at https://github.com/cspzyy/RealHiTBench.

📄 PDF Abstract BibTeX arXiv:2506.13405

Code (1)

cspzyy/realhitbench 공식 구현

Methods 이 논문이 사용한 방법론

Focus 설명 없음

Similar Papers 제목 키워드 기반

MPDocBench-Parse: Benchmarking Practical Multi-page Document Parsing

2026-05-21 · Bangbang Zhou, Hangdi Xing, Yifan Chen, Jianjun Xu 외 arxiv

Document parsing converts visually rich documents into machine-readable structured representations, forming a crucial foundation for information systems. Although many benchmarks have been proposed for document parsing, …

GUI-CEval: A Hierarchical and Comprehensive Chinese Benchmark for Mobile GUI Agents

2026-03-16 · Yang Li, Yuchen Liu, Haoyu Lu, Zhiqiang Xia 외 arxiv

Recent progress in Multimodal Large Language Models (MLLMs) has enabled mobile GUI agents capable of visual perception, cross-modal reasoning, and interactive control. However, existing benchmarks are largely English-cen…

AGIDefect-4K: A Richly Annotated Dataset for AI-Generated Image Defect Detection, Localization and Explanation

2026-08-21 · Xiangfei Sheng, Weidong Zou, Tianjiao Gu, Zhichao Yang 외 arxiv

Generative AI can now produce highly realistic images, yet current models still exhibit subtle but critical defects that undermine their reliability. While existing AI-generated image (AGI) evaluation benchmarks have mad…

OmniReview: A Large-scale Benchmark and LLM-enhanced Framework for Realistic Reviewer Recommendation

2026-02-09 · Yehua Huang, Penglei Sun, Zebin Chen, Zhenheng Tang 외 arxiv

Academic peer review remains the cornerstone of scholarly validation, yet the field faces some challenges in data and methods. From the data perspective, existing research is hindered by the scarcity of large-scale, veri…

Multi-Task Learning

MultiHiertt: Numerical Reasoning over Multi Hierarchical Tabular and Textual Data

2022-06-03 · ACL 2022 5 · Yilun Zhao, Yunxiang Li, Chenying Li, Rui Zhang

Numerical reasoning over hybrid data containing both textual and tabular content (e.g., financial reports) has recently attracted much attention in the NLP community. However, existing question answering (QA) benchmarks …

Question Answering