paper-with-me

홈 › Papers

Cost of Structural Learning Under Censored Feedback: A Threshold-Bandit Approach

2026-05-26 · Michael Ledford, William Regli arxiv

In many multi-agent applications, tasks yield rewards only when executed by a coalition meeting an unknown size threshold; otherwise, feedback is fully censored. This censorship creates an identifiability problem: agents cannot distinguish stochastic failure from insufficient coordination. We formalize this setting as the Threshold-Activated Cooperative Multi-Armed Bandit (TAC-MAB) and analyze it under both centralized and decentralized coordination. We show that a centralized algorithm (C-TAC) achieves cumulative regret O(log T), decomposed into a structural-search term that captures the cost of resolving feasibility under censored feedback and a statistical-monitoring term for value estimation. We then introduce D-TAC, a decentralized event-triggered protocol in which agents synchronize only when their structural beliefs change. Empirically, D-TAC achieves a 23x reduction in communication relative to the centralized baseline while preserving feasibility alignment under conservative belief fusion. These results characterize the coordination cost of learning under censored feedback and show that near-centralized communication efficiency is achievable without continuous synchronization.

📄 PDF Abstract BibTeX arXiv:2605.27076

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Threshold Bandits, With and Without Censored Feedback

2016-12-01 · NeurIPS 2016 12 · Jacob D. Abernethy, Kareem Amin, Ruihao Zhu

We consider the \emph{Threshold Bandit} setting, a variant of the classical multi-armed bandit problem in which the reward on each round depends on a piece of side information known as a \emph{threshold value}. The learn…

Community detection in censored hypergraph

2021-11-04 · Mingao Yuan, Bin Zhao, Xiaofeng Zhao

Community detection refers to the problem of clustering the nodes of a network (either graph or hypergrah) into groups. Various algorithms are available for community detection and all these methods apply to uncensored n…

Community DetectionMissing Values

Generalization Error Bounds for Learning under Censored Feedback

2024-04-14 · Yifan Yang, Ali Payani, Parinaz Naghizadeh

Generalization error bounds from learning theory provide statistical guarantees on how well an algorithm will perform on previously unseen data. In this paper, we characterize the impacts of data non-IIDness due to censo…

Learning TheoryRecommendation Systems

Censored Semi-Bandits: A Framework for Resource Allocation with Censored Feedback

2019-09-04 · NeurIPS 2019 12 · Arun Verma, Manjesh K. Hanawal, Arun Rajkumar, Raman Sankaran

In this paper, we study censored Semi-Bandits, a novel variant of the semi-bandits problem. The learner is assumed to have a fixed amount of resources, which it allocates to the arms at each time step. The loss observed …

Multi-Armed Bandits

Data Driven Block Replacement Scheduling

2026-07-16 · Aniruddhan Ganesaraman, VIdyadhar Kulkarni arxiv

We develop data-driven algorithms for maintaining $N$ independent identical machines under a \textit{block replacement policy}, in which each machine is replaced upon failure and all machines are jointly replaced at regu…