paper-with-me

홈 › Papers

Tree-Guided Identify-Then-Exploit: A Unified Framework of Best Arm Identification and Regret Minimization for Dueling Bandits

2026-06-01 · Pu Wang, Yao-Xiang Ding arxiv

We study $N$-armed stochastic dueling bandits under the Condorcet-winner assumption, where three widely adopted objectives are considered: best-arm identification (BAI), weak regret, and strong regret. We propose Tree-Guided Identify-Then-Exploit (TG-ITE), the first unified framework to tackle all these objectives to our knowledge. Without requiring stronger assumptions, we propose a shared tree-guided identification approach to find a high-confidence incumbent within $O(N)$ comparisons. We further propose varied exploitation strategies to utilize this warm-start stage to optimize the specific objectives at hand. This methodology enables our approach to (1) achieve $O(N)$ sample complexity in BAI without commonly adopted stronger assumptions; (2) build the first winner-stays-style algorithm to achieve $O(N)$ weak regret; (3) enjoy the same $O(N \log T)$ guarantee as specialized strong-regret approaches; (4) realize the joint optimization of BAI and weak regret with $O(N)$ guarantees for both, eliminating the sub-optimal gap of $O(\log N)$ in the existing approach. Our results provide evidence that the trade-off between BAI and regret minimization is relatively benign in dueling bandits.

📄 PDF Abstract BibTeX arXiv:2606.01799

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Spend Search Where It Pays: Value-Guided Structured Sampling and Optimization for Generative Recommendation

2026-02-11 · Jie Jiang, Yangru Huang, Zeyu Wang, Changping Wang 외 arxiv

Generative recommendation via autoregressive models has unified retrieval and ranking into a single conditional generation framework. However, fine-tuning these models with Reinforcement Learning (RL) often suffers from …

Reinforcement Learning

DIAL-GS: Dynamic Instance Aware Reconstruction for Label-free Street Scenes with 4D Gaussian Splatting

2025-11-10 · Chenpeng Su, Wenhua Wu, Chensheng Peng, Tianchen Deng 외 arxiv

Urban scene reconstruction is critical for autonomous driving, enabling structured 3D representations for data synthesis and closed-loop testing. Supervised approaches rely on costly human annotations and lack scalabilit…

Autonomous Driving

mTREE: Multi-Level Text-Guided Representation End-to-End Learning for Whole Slide Image Analysis

2024-05-28 · Quan Liu, Ruining Deng, Can Cui, Tianyuan Yao 외

Multi-modal learning adeptly integrates visual and textual data, but its application to histopathology image and text analysis remains challenging, particularly with large, high-resolution images like gigapixel Whole Sli…

Survival Predictionwhole slide images

EvalTree: Profiling Language Model Weaknesses via Hierarchical Capability Trees

2025-03-11 · Zhiyuan Zeng, Yizhong Wang, Hannaneh Hajishirzi, Pang Wei Koh

An ideal model evaluation should achieve two goals: identifying where the model fails and providing actionable improvement guidance. Toward these goals for Language Model (LM) evaluations, we formulate the problem of gen…

ChatbotLanguage ModelingLanguage ModellingMath

UDWiki: guided creation and exploitation of UD treebanks

2021-12-01 · UDW (SyntaxFest) 2021 12 · Maarten Janssen