paper-with-me

홈 › Papers

EvalTree: Profiling Language Model Weaknesses via Hierarchical Capability Trees

2025-03-11 · Zhiyuan Zeng, Yizhong Wang, Hannaneh Hajishirzi, Pang Wei Koh

An ideal model evaluation should achieve two goals: identifying where the model fails and providing actionable improvement guidance. Toward these goals for Language Model (LM) evaluations, we formulate the problem of generating a weakness profile, a set of weaknesses expressed in natural language, given an LM's performance on every individual instance in a benchmark. We introduce a suite of quantitative assessments to compare different weakness profiling methods. We also propose a weakness profiling method EvalTree. It constructs a capability tree where each node represents a capability described in natural language and is linked to a subset of benchmark instances that specifically evaluate this capability; it then extracts nodes where the LM performs poorly to generate a weakness profile. On the MATH and WildChat benchmarks, we show that EvalTree outperforms baseline weakness profiling methods by identifying weaknesses more precisely and comprehensively. Weakness profiling further enables weakness-guided data collection, and training data collection guided by EvalTree-identified weaknesses improves LM performance more than other data collection strategies. We also show how EvalTree exposes flaws in Chatbot Arena's human-voter-based evaluation practice. To facilitate future work, we release our code and an interface that allows practitioners to interactively explore the capability trees built by EvalTree.

📄 PDF Abstract BibTeX arXiv:2503.08893

Code (1)

zhiyuan-zeng/evaltree 공식 구현

Tasks

ChatbotLanguage ModelingLanguage ModellingMath

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically

Similar Papers 제목 키워드 기반

CRAFT: Clustering Rubrics to Diagnose Weak LLM Capabilities and Generate Targeted Fine-Tuning Data

2026-07-17 · Vipul Gupta, Zihao Wang, Razvan-Gabriel Dumitru, MohammadHossein Rezaei 외 arxiv

Evaluations should do more than measure a models current performance. They should tell us what to fix for the next model iteration and provide a way to generate targeted post training data. Most evaluation pipelines iden…

EBench: Elemental Diagnosis of Generalist Mobile Manipulation Policies

2026-06-16 · Ning Gao, Jinliang Zheng, Xing Gao, Haoxiang Ma 외 arxiv

We present EBench, a simulation benchmark that diagnoses generalist mobile manipulation policies beyond a single success-rate scalar. EBench comprises 26 diverse and challenging manipulation tasks annotated along 5 capab…

Evaluating Character Understanding of Large Language Models via Character Profiling from Fictional Works

2024-04-19 · Xinfeng Yuan, Siyu Yuan, Yuhan Cui, Tianhe Lin 외

Large language models (LLMs) have demonstrated impressive performance and spurred numerous AI applications, in which role-playing agents (RPAs) are particularly popular, especially for fictional characters. The prerequis…

A Usage-centric Take on Intent Understanding in E-Commerce

2024-02-22 · Wendi Zhou, Tianyi Li, Pavlos Vougiouklis, Mark Steedman 외

Identifying and understanding user intents is a pivotal task for E-Commerce. Despite its essential role in product recommendation and business user profiling analysis, intent understanding has not been consistently defin…

Product Recommendation

Effectiveness of LLMs in Temporal User Profiling for Recommendation

2025-10-31 · Milad Sabouri, Masoud Mansoury, Kun Lin, Bamshad Mobasher arxiv

Effectively modeling the dynamic nature of user preferences is crucial for enhancing recommendation accuracy and fostering transparency in recommender systems. Traditional user profiling often overlooks the distinction b…