paper-with-me

홈 › Papers

UniRank: Unimodal Bandit Algorithm for Online Ranking

2022-08-02 · Camille-Sovanneary Gauthier, Romaric Gaudel, Elisa Fromont

We tackle a new emerging problem, which is finding an optimal monopartite matching in a weighted graph. The semi-bandit version, where a full matching is sampled at each iteration, has been addressed by \cite{ADMA}, creating an algorithm with an expected regret matching $O(\frac{L\log(L)}{\Delta}\log(T))$ with $2L$ players, $T$ iterations and a minimum reward gap $\Delta$. We reduce this bound in two steps. First, as in \cite{GRAB} and \cite{UniRank} we use the unimodality property of the expected reward on the appropriate graph to design an algorithm with a regret in $O(L\frac{1}{\Delta}\log(T))$. Secondly, we show that by moving the focus towards the main question `\emph{Is user $i$ better than user $j$?}' this regret becomes $O(L\frac{\Delta}{\tilde{\Delta}^2}\log(T))$, where $\Tilde{\Delta} > \Delta$ derives from a better way of comparing users. Some experimental results finally show these theoretical results are corroborated in practice.

📄 PDF Abstract BibTeX arXiv:2208.01515

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Unimodal Mono-Partite Matching in a Bandit Setting

2022-08-02 · Romaric Gaudel, Matthieu Rodet

We tackle a new emerging problem, which is finding an optimal monopartite matching in a weighted graph. The semi-bandit version, where a full matching is sampled at each iteration, has been addressed by \cite{ADMA}, crea…

UniRank: End-to-End Domain-Specific Reranking of Hybrid Text-Image Candidates

2026-02-08 · Yupei Yang, Lin Yang, Wanxi Deng, Lin Qu 외 arxiv

Reranking is a critical component in many information retrieval pipelines. Despite remarkable progress in text-only settings, multimodal reranking remains challenging, particularly when the candidate set contains hybrid …

Reinforcement LearningInformation RetrievalDomain Adaptation

UniRank: A Multi-Agent Calibration Pipeline for Estimating University Rankings from Anonymized Bibliometric Signals

2026-02-21 · Pedram Riyazimehr, Seyyed Ehsan Mahmoudi arxiv

We present UniRank, a multi-agent LLM pipeline that estimates university positions across global ranking systems using only publicly available bibliometric data from OpenAlex and Semantic Scholar. The system employs a th…

Optimizing Ranking Systems Online as Bandits

2021-10-12 · Chang Li

Ranking system is the core part of modern retrieval and recommender systems, where the goal is to rank candidate items given user contexts. Optimizing ranking systems online means that the deployed system can serve user …

Learning-To-RankOnline Ranker EvaluationRecommendation SystemsRetrieval

Thompson Sampling for Unimodal Bandits

2021-06-15 · Long Yang, Zhao Li, Zehong Hu, Shasha Ruan 외

In this paper, we propose a Thompson Sampling algorithm for \emph{unimodal} bandits, where the expected reward is unimodal over the partially ordered arms. To exploit the unimodal structure better, at each step, instead …

Thompson Sampling