A Fisher's exact test justification of the TF-IDF term-weighting scheme
Term frequency-inverse document frequency, or TF-IDF for short, is arguably the most celebrated mathematical expression in the history of information retrieval. Conceived as a simple heuristic quantifying the extent to which a given term's occurrences are concentrated in any one given document out of many, TF-IDF and its many variants are routinely used as term-weighting schemes in diverse text analysis applications. There is a growing body of scholarship dedicated to placing TF-IDF on a sound theoretical foundation. Building on that tradition, this paper justifies the use of TF-IDF to the statistics community by demonstrating how the famed expression can be understood from a significance testing perspective. We show that the common TF-IDF variant TF-ICF is, under mild regularity conditions, closely related to the negative logarithm of the $p$-value from a one-tailed version of Fisher's exact test of statistical significance. As a corollary, we establish a connection between TF-IDF and the said negative log-transformed $p$-value under certain idealized assumptions. We further demonstrate, as a limiting case, that this same quantity converges to TF-IDF in the limit of an infinitely large document collection. The Fisher's exact test justification of TF-IDF equips the working statistician with a ready explanation of the term-weighting scheme's long-established effectiveness.
Code (0)
등록된 구현이 없습니다.
Tasks
Information RetrievalSimilar Papers 제목 키워드 기반
\'Evaluation de mesures d'association pour les bigrammes et les trigrammes au moyen du test exact de Fisher (Using Fisher's Exact Test to Evaluate Association Measures for Bigrams and Trigrams)
Pour d{\'e}terminer si certaines mesures d{'}association lexicale fr{\'e}quemment employ{\'e}es en TAL attribuent des scores {\'e}lev{\'e}s {\`a} des n-grammes que le hasard aurait pu produire aussi souvent qu{'}observ{\…
es-enR. A. Fisher's Exact Test Revisited
This note provides a conceptual clarification of Ronald Aylmer Fisher's (1935) pioneering exact test in the context of the Lady Testing Tea experiment. It unveils a critical implicit assumption in Fisher's calibration: t…
Discussion of `Multiscale Fisher's Independence Test for Multivariate Dependence'
We discuss how MultiFIT, the Multiscale Fisher's Independence Test for Multivariate Dependence proposed by Gorsky and Ma (2022), compares to existing linear-time kernel tests based on the Hilbert-Schmidt independence cri…
Using Fisher's Exact Test to Evaluate Association Measures for N-grams
To determine whether some often-used lexical association measures assign high scores to n-grams that chance could have produced as frequently as observed, we used an extension of Fisher's exact test to sequences longer t…
Exact path-integral representation of the Wright-Fisher model with mutation and selection
The Wright-Fisher model describes a biological population containing a finite number of individuals. In this work we consider a Wright-Fisher model for a randomly mating population, where selection and mutation act at an…