paper-with-me

홈 › Papers

Common TF-IDF variants arise as key components in the test statistic of a penalized likelihood-ratio test for word burstiness

2026-04-01 · Zeyad Ahmed, Paul Sheridan, Michael McIsaac, Aitazaz A. Farooque arxiv

TF-IDF is a classical formula that is widely used for identifying important terms within documents. We show that TF-IDF-like scores arise naturally from the test statistic of a penalized likelihood-ratio test setup capturing word burstiness (also known as word over-dispersion). In our framework, the alternative hypothesis captures word burstiness by modeling a collection of documents according to a family of beta-binomial distributions with a gamma penalty term on the precision parameter. In contrast, the null hypothesis assumes that words are binomially distributed in collection documents, a modeling approach that fails to account for word burstiness. We find that a term-weighting scheme given rise to by this test statistic performs comparably to TF-IDF on document classification tasks. This paper provides insights into TF-IDF from a statistical perspective and underscores the potential of hypothesis testing frameworks for advancing term-weighting scheme development.

📄 PDF Abstract BibTeX arXiv:2604.00672

Code (0)

등록된 구현이 없습니다.

Tasks

Document Classification

Similar Papers 제목 키워드 기반

Towards Learning and Verifying Invariants of Cyber-Physical Systems by Code Mutation

2016-09-06 · Yuqi Chen, Christopher M. Poskitt, Jun Sun

Cyber-physical systems (CPS), which integrate algorithmic control with physical processes, often consist of physically distributed components communicating over a network. A malfunctioning or compromised component in suc…

Statistics of shared components in complex component systems

2018-04-23

Many complex systems are modular. Such systems can be represented as "component systems", i.e., sets of elementary components, such as LEGO bricks in LEGO sets. The bricks found in a LEGO set reflect a target architectur…

Robust Joint and Individual Variance Explained

2017-07-01 · CVPR 2017 7 · Christos Sagonas, Yannis Panagakis, Alina Leidinger, Stefanos Zafeiriou

Discovering the common (joint) and individual subspaces is crucial for analysis of multiple data sets, including multi-view and multi-modal data. Several statistical machine learning methods have been developed for disco…

Adaptivity and Computation-Statistics Tradeoffs for Kernel and Distance based High Dimensional Two Sample Testing

2015-08-04 · Aaditya Ramdas, Sashank J. Reddi, Barnabas Poczos, Aarti Singh 외

Nonparametric two sample testing is a decision theoretic problem that involves identifying differences between two random variables without making parametric assumptions about their underlying distributions. We refer to …

Two-sample testing

Functional Analysis of Variance for Association Studies

2025-08-14 · Olga A. Vsevolozhskaya, Dmitri V. Zaykin, Mark C. Greenwood, Changshuai Wei 외 arxiv

While progress has been made in identifying common genetic variants associated with human diseases, for most of common complex diseases, the identified genetic variants only account for a small proportion of heritability…