paper-with-me

홈 › Papers

Metadata Predictability Is Not Evidence Dependence: An Intervention-Based Audit for Weak-Label Benchmarks

2026-05-22 · Kan Shao arxiv

We study a protocol-level test for weak-label benchmarks: whether benchmark outputs change when the provided evidence is intervened on. Metadata-only shortcut checks answer a different question, namely whether outputs are predictable from metadata priors. We therefore combine a metadata statistic, the Metadata Prior Dominance Score (MPDS), with an evidence-intervention statistic, ΔEvi, measuring sensitivity to evidence identity under cross-item shuffling. Synthetic HotpotQA gives a constructed counterexample to metadata-only screening: MPDS is only moderate (0.643), yet ΔEvi is zero. Stronger-reader reruns show why calibration belongs in the test procedure: SNLI shows a calibration reversal, reconstructed HotpotQA occupies a question-dominant warning region, and FEVER is a strongly evidence-sensitive positive control across four transformers. The practical lesson is simple: benchmark audits should report metadata-only screening, evidence intervention, and reader-strength calibration together.

📄 PDF Abstract BibTeX arXiv:2605.23701

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Spurious Predictability in Financial Machine Learning

2026-04-16 · Sotirios D. Nikolopoulos arxiv

Adaptive specification search generates statistically significant backtests even under martingale-difference nulls. We introduce a falsification audit testing complete predictive workflows against synthetic reference cla…

The World is Not Mono: Enabling Spatial Understanding in Large Audio-Language Models

2026-01-06 · Yuhuan You, Lai Wei, Xihong Wu, Tianshu Qu arxiv

Large audio-language models have made rapid progress in recognizing what is present in an audio clip, but spatial audio-language understanding still lacks a clear task interface. A model must also decide where sound even…

DoAtlas-1: A Causal Compilation Paradigm for Clinical AI

2026-02-22 · Yulong Li, Jianxu Chen, Xiwei Liu, Chuanyue Suo 외 arxiv

Medical foundation models generate narrative explanations but cannot quantify intervention effects, detect evidence conflicts, or validate literature claims, limiting clinical auditability. We propose causal compilation,…

Text Generation

ReplaySCM: A Benchmark for Executable Causal Mechanism Induction from Interventions

2026-05-05 · Serafim Batzoglou arxiv

Most causal benchmarks for language models score local answers or graph structure. We introduce ReplaySCM, a 1,300 item benchmark for executable causal mechanism induction from finite interventional evidence. Each item c…

Quantile Dependence between Stock Markets and its Application in Volatility Forecasting

2016-08-25

This paper examines quantile dependence between international stock markets and evaluates its use for improving volatility forecasting. First, we analyze quantile dependence and directional predictability between the US …