paper-with-me

홈 › Papers

Quantifying Ranking Instability Across Evaluation Protocol Axes in Gene Regulatory Network Benchmarking

2026-03-03 · Ihor Kendiukhov arxiv

Benchmark rankings are routinely used to justify scientific claims about method quality in gene regulatory network (GRN) inference, yet the stability of these rankings under plausible evaluation protocol choices is rarely examined. We present a systematic diagnostic framework for measuring ranking instability under protocol shift, including decomposition tools that separate base rate effects from discrimination effects. Using existing single cell GRN benchmark outputs across three human tissues and six inference methods, we quantify pairwise reversal rates across four protocol axes: candidate set restriction (16.3 percent, 95 percent CI 11.0 to 23.4 percent), tissue context (19.3 percent), reference network choice (32.1 percent), and symbol mapping policy (0.0 percent). A permutation null confirms that observed reversal rates are far below random order expectations (0.163 versus null mean 0.500), indicating partially stable but non invariant ranking structure. Our decomposition reveals that reversals are driven by changes in the relative discrimination ability of methods rather than by base rate inflation, a finding that challenges a common implicit assumption in GRN benchmarking. We propose concrete reporting practices for stability aware evaluation and provide a diagnostic toolkit for identifying method pairs at risk of reversal.

📄 PDF Abstract BibTeX arXiv:2603.03493

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Beyond the Fold: Quantifying Split-Level Noise and the Case for Leave-One-Dataset-Out AU Evaluation

2026-04-02 · Saurabh Hinduja, Gurmeet Kaur, Maneesh Bilalpur, Jeffrey Cohn 외 arxiv

Subject-exclusive cross-validation is the standard evaluation protocol for facial Action Unit (AU) detection, yet reported improvements are often small. We show that cross-validation itself introduces measurable stochast…

MaP: A Unified Framework for Reliable Evaluation of Pre-training Dynamics

2025-10-10 · Jiapeng Wang, Changxin Tian, Kunlong Chen, Ziqi Liu 외 arxiv

Reliable evaluation is fundamental to the progress of Large Language Models (LLMs), yet the evaluation process during pre-training is plagued by significant instability that obscures true learning dynamics. In this work,…

Quantifying Consistency in LLM Logical Reasoning via Structural Uncertainty

2026-06-15 · Baishali Chaudhury, Mengdie Flora Wang, Hyunji Hayley Park, Rahul Ghosh 외 arxiv

Large language models can arrive at the same answer through reasoning paths that are unstable, contradictory, or difficult to rank consistently -- a failure mode especially prevalent in multi-step deductive reasoning. Ex…

Mathematical ReasoningLogical Reasoning

Tempora: Characterising the Time-Contingent Utility of Online Test-Time Adaptation

2026-02-05 · Sudarshan Sreeram, Young D. Kwon, Cecilia Mascolo arxiv

Test-time adaptation (TTA) offers a compelling remedy for machine learning (ML) models that degrade under domain shifts, improving generalisation on-the-fly with only unlabelled samples. This flexibility suits real deplo…

Test-time Adaptation

Cross-Lingual Stability of LLM Judges Under Controlled Generation: Evidence from Finno-Ugric Languages

2026-02-02 · Isaac Chung, Linda Freienthal arxiv

Cross-lingual evaluation of large language models (LLMs) typically conflates two sources of variance: genuine model performance differences and measurement instability. We investigate evaluation reliability by holding ge…

Semantic Similarity