paper-with-me

홈 › Papers

TabXEval: Why this is a Bad Table? An eXhaustive Rubric for Table Evaluation

2025-05-28 · Vihang Pancholi, Jainit Bafna, Tejas Anvekar, Manish Shrivastava, Vivek Gupta

Evaluating tables qualitatively & quantitatively presents a significant challenge, as traditional metrics often fail to capture nuanced structural and content discrepancies. To address this, we introduce a novel, methodical rubric integrating multi-level structural descriptors with fine-grained contextual quantification, thereby establishing a robust foundation for comprehensive table comparison. Building on this foundation, we propose TabXEval, an eXhaustive and eXplainable two-phase evaluation framework. TabXEval initially aligns reference tables structurally via TabAlign & subsequently conducts a systematic semantic and syntactic comparison using TabCompare; this approach clarifies the evaluation process and pinpoints subtle discrepancies overlooked by conventional methods. The efficacy of this framework is assessed using TabXBench, a novel, diverse, multi-domain benchmark we developed, featuring realistic table perturbations and human-annotated assessments. Finally, a systematic analysis of existing evaluation methods through sensitivity-specificity trade-offs demonstrates the qualitative and quantitative effectiveness of TabXEval across diverse table-related tasks and domains, paving the way for future innovations in explainable table evaluation.

📄 PDF Abstract BibTeX arXiv:2505.22176

Code (0)

등록된 구현이 없습니다.

Tasks

Specificity

Similar Papers 제목 키워드 기반

ExecRubrics: Executable Tool-Augmented Rubrics for Verifiable and Efficient Long-Form Evaluation

2026-08-23 · Kaustubh D. Dhole, Charles L. A. Clarke, Eugene Y. Agichtein arxiv

Rubrics aim to make language-model evaluation transparent by decomposing response quality into interpretable criteria. However, natural-language rubrics are often ambiguous, require LLM judges, and typically assume crite…

RubricRAG: Towards Interpretable and Reliable LLM Evaluation via Domain Knowledge Retrieval for Rubric Generation

2026-03-21 · Kaustubh D. Dhole, Eugene Agichtein arxiv

Large language models (LLMs) are increasingly evaluated and sometimes trained using automated graders such as LLM-as-judges that output scalar scores or preferences. While convenient, these approaches are often opaque: a…

RewardVerse: Rubric-Guided Policy Optimization for Video Reward Modeling

2026-09-19 · Zhenchen Tang, Yang Li, Songlin Yang, Bo Peng 외 hf

Reinforcement learning (RL) is vital for optimizing video generation models, with a robust reward model (RM) serving as the cornerstone. However, existing video reward models often produce unstable scalar scores because …

Reinforcement LearningVideo Generation

From Rubrics to Reliable Scores: Evidence-Grounded Text Evaluation with LLM Judges

2026-01-13 · Yihan Hong, Huaiyuan Yao, Bolin Shen, Wanpeng Xu 외 arxiv

Rubric-based text evaluation increasingly relies on large language models (LLMs) as scalable judges, yet frozen black-box models can interpret the same criteria inconsistently, produce score attributions that are difficu…

Text Generation

AMARIS: A Memory-Augmented Rubric Improvement System for Rubric-Based Reinforcement Learning

2026-05-18 · Peilin Wu, Xinlu Zhang, Kun Wan, Wentian Zhao 외 arxiv

Rubric-based reward shaping provides interpretable and editable reward signals for fine-tuning LLMs via reinforcement learning (RL), but existing adaptive rubric methods typically update criteria from local evidence such…

Reinforcement LearningInstruction Following