paper-with-me

홈 › Papers

ReLE: A Scalable System and Structured Benchmark for Diagnosing Capability Anisotropy in Chinese LLMs

2026-01-24 · Rui Fang, Jian Li, Wei Chen, Bin Hu, Ying-Cong Chen, Xin Tang, Liang Diao arxiv

Large Language Models (LLMs) have achieved rapid progress in Chinese language understanding, yet accurately evaluating their capabilities remains challenged by benchmark saturation and prohibitive computational costs. While static leaderboards provide snapshot rankings, they often mask the structural trade-offs between capabilities. In this work, we present ReLE (Robust Efficient Live Evaluation), a scalable system designed to diagnose Capability Anisotropy, the non-uniformity of model performance across domains. Using ReLE, we evaluate 304 models (189 commercial, 115 open-source) across a Domain $\times$ Capability orthogonal matrix comprising 207,843 samples. We introduce two methodological contributions to address current evaluation pitfalls: (1) A Symbolic-Grounded Hybrid Scoring Mechanism that eliminates embedding-based false positives in reasoning tasks; (2) A Dynamic Variance-Aware Scheduler based on Neyman allocation with noise correction, which reduces compute costs by 70\% compared to full-pass evaluations while maintaining a ranking correlation of $ρ=0.96$. Our analysis reveals that aggregate rankings are highly sensitive to weighting schemes: models exhibit a Rank Stability Amplitude (RSA) of 11.4 in ReLE versus $\sim$5.0 in traditional benchmarks, confirming that modern models are highly specialized rather than generally superior. We position ReLE not as a replacement for comprehensive static benchmarks, but as a high-frequency diagnostic monitor for the evolving model landscape.

📄 PDF Abstract BibTeX arXiv:2601.17399

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Reasoning Structure of Large Language Models

2026-06-02 · Frédéric Berdoz, Luca A. Lanzendörfer, Fabian Farestam, Roger Wattenhofer arxiv

Large reasoning models (LRMs) are often evaluated using metrics such as final-answer accuracy or token count. However, identical scores on these metrics can hide fundamentally different reasoning structures. To address t…

Diagnosing Causal Reasoning in Vision-Language Models via Structured Relevance Graphs

2026-02-24 · Dhita Putri Pratama, Soyeon Caren Han, Yihao Ding arxiv

Large Vision-Language Models (LVLMs) achieve strong performance on visual question answering benchmarks, yet often rely on spurious correlations rather than genuine causal reasoning. Existing evaluations primarily assess…

Visual Question AnsweringCausal Inference

MedEinst: Benchmarking the Einstellung Effect in Medical LLMs through Counterfactual Differential Diagnosis

2026-01-10 · Wenting Chen, Zhongrui Zhu, Guolin Huang, Wenxuan Wang arxiv

Despite achieving high accuracy on medical benchmarks, LLMs exhibit the Einstellung Effect in clinical diagnosis--relying on statistical shortcuts rather than patient-specific evidence, causing misdiagnosis in atypical c…

Causal Inference

Diagnosing and Addressing Pitfalls in KG-RAG Datasets: Toward More Reliable Benchmarking

2025-05-29 · Liangliang Zhang, Zhuorui Jiang, Hongliang Chi, Haoyang Chen 외

Knowledge Graph Question Answering (KGQA) systems rely on high-quality benchmarks to evaluate complex multi-hop reasoning. However, despite their widespread use, popular datasets such as WebQSP and CWQ suffer from critic…

BenchmarkingGraph Question AnsweringQuestion AnsweringRAG

LogSieve: Task-Aware CI Log Reduction for Sustainable LLM-Based Analysis

2026-01-28 · Marcus Emmanuel Barnes, Taher A. Ghaleb, Safwat Hassan arxiv

Logs are essential for understanding Continuous Integration (CI) behavior, particularly for diagnosing build failures and performance regressions. Yet their growing volume and verbosity make both manual inspection and au…

Anomaly Detection