paper-with-me

Papers

Bradley-Terry Rankings for Recommender Systems Across Dataset Taxonomies

2026-06-05 · Ekaterina Grishina, Stepan Kuznetsov, Askar Tsyganov, Ilya Ivanov, Daria Korovaitceva, Margarita Rusanova, Uliana Parkina, Alexander Derevyagin, Evgeny Frolov, Sergey Samsonov, Anton Lysenko arxiv

The ranking of recommendation algorithms is a challenging problem since model performance is sensitive to dataset characteristics such as sparsity, sequential structure, and scale. This drives a demand for a proper methodology for fair comparison between algorithms. Naive aggregation of performance metrics (e.g., averaging NDCG over benchmarks) can yield misleading rankings, undermining practical selection. To address this problem, we introduce a novel, data-driven ranking methodology based on Bradley-Terry (BT) model. We demonstrate that the obtained ranking depends on key dataset statistics. Additionally, we propose a novel metric for evaluating ranking consistency and demonstrate robustness of our ranking to incomplete data. Finally, we introduce a dataset-specific methodology for ranking algorithms on unseen datasets without running the models, relying on extensions of the Bradley-Terry framework, including BT trees and BT models with covariates.

📄 PDF Abstract BibTeX arXiv:2606.07492

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Nonparametric Estimation in the Dynamic Bradley-Terry Model

2020-02-28 · Heejong Bong, Wanshan Li, Shamindra Shrotriya, Alessandro Rinaldo

We propose a time-varying generalization of the Bradley-Terry model that allows for nonparametric modeling of dynamic global rankings of distinct teams. We develop a novel estimator that relies on kernel smoothing to pre…

model

Personalized Benchmarking: Evaluating LLMs by Individual Preferences

2026-04-21 · Cristina Garbacea, Heran Wang, Chenhao Tan arxiv

With the rise in capabilities of large language models (LLMs) and their deployment in real-world tasks, evaluating LLM alignment with human preferences has become an important challenge. Current benchmarks average prefer…

On rankings in multiplayer games with an application to the game of Whist

2026-04-01 · Alexis Coyette, Charles Modera, Candy Sonveaux, Judicaël Mohet 외 arxiv

We propose a novel extension of the Bradley-Terry model to multiplayer games and adapt a recent algorithm by Newman [1] to our model. We demonstrate the use of our proposed method on synthetic datasets and on a real data…

Efficient Bayesian Inference from Noisy Pairwise Comparisons

2025-10-10 · Till Aczel, Lucas Theis, Roger Wattenhofer arxiv

Evaluating generative models is challenging because standard metrics often fail to reflect human preferences. Human evaluations are more reliable but costly and noisy, as participants vary in expertise, attention, and di…

Bayesian Inference

Ranking and Selection from Pairwise Comparisons: Empirical Bayes Methods for Citation Analysis

2021-12-21 · Jiaying Gu, Roger Koenker

We study the Stigler model of citation flows among journals adapting the pairwise comparison model of Bradley and Terry to do ranking and selection of journal influence based on nonparametric empirical Bayes procedures. …