The evidence contained in the P-value is context dependent
In a recent opinion article, Muff et al. recapitulate well-known objections to the Neyman-Pearson Null-Hypothesis Significance Testing (NHST) framework and call for reforming our practices in statistical reporting. We agree with them on several important points: the significance threshold P<0.05 is only a convention, chosen as a compromise between type I and II error rates; transforming the p-value into a dichotomous statement leads to a loss of information; and p-values should be interpreted together with other statistical indicators, in particular effect sizes and their uncertainty. In our view, a lot of progress in reporting results can already be achieved by keeping these three points in mind. We were surprised and worried, however, by Muff et al.'s suggestion to interpret the p-value as a "gradual notion of evidence". Muff et al. recommend, for example, that a P-value > 0.1 should be reported as "little or no evidence" and a P-value of 0.001 as "strong evidence" in favor of the alternative hypothesis H1.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
In-Context Learning Without Copying
Induction heads are attention heads that perform inductive copying by matching patterns from earlier context and copying their continuations verbatim. As models develop induction heads, they experience a sharp drop in tr…
Adaptive Evidence Weighting for Audio-Spatiotemporal Fusion
Many machine learning systems have access to multiple sources of evidence for the same prediction target, yet these sources often differ in reliability and informativeness across inputs. In bioacoustic classification, sp…
Bayesian InferenceCORRECT: Context- and Reference-Augmented Reasoning and Prompting for Fact-Checking
Fact-checking the truthfulness of claims usually requires reasoning over multiple evidence sentences. Oftentimes, evidence sentences may not be always self-contained, and may require additional contexts and references fr…
Fact CheckingSelf-Paced Context Evaluation for Contextual Reinforcement Learning
Reinforcement learning (RL) has made a lot of advances for solving a single problem in a given environment; but learning policies that generalize to unseen variations of a problem remains challenging. To improve sample e…
reinforcement-learningReinforcement LearningReinforcement Learning (RL)Self-contained Beta-with-Spikes Approximation for Inference Under a Wright-Fisher Model
We construct a reliable estimation of evolutionary parameters within the Wright-Fisher model, which describes changes in allele frequencies due to selection and genetic drift, from time-series data. Such data exists for …
Time SeriesTime Series Analysis