paper-with-me

홈 › Papers

Short-Term Pain for Long-Term Gain: Adaptive Experiment with Post-Commitment Reward Shift

2026-07-26 · Puping Jiang, Wei Tang arxiv

Decision-makers in learning environments face a dilemma when their short-term optimal actions may not favor their long-term benefits the most. To understand the fundamental tradeoff behind the dilemma, we study adaptive experimentation with post-commitment reward shifts. During an experiment phase, the decision-maker may adaptively test multiple options; during a subsequent commitment phase, the decision-maker must commit to a single option, whose reward may differ from its pre-commitment reward. We propose the Reserved Arm Eliminations for Commitment (RAEC) algorithm, which reserves a predetermined portion of the experiment phase to identify the best post-shift option while using the remaining rounds to minimize short-run regret. We establish regret upper bounds for RAEC across all parameter regimes and matching minimax lower bounds, providing a tight characterization of the cost of balancing short-term performance and long-term commitment. We also study two extensions. With prior structural knowledge linking pre- and post-shift rewards, we show that correctly identifying the ranking-changing component of the shift is more important than estimating its absolute magnitude. For settings with concave commitment rewards and portfolio choice, we develop the Reserved Online Stochastic Convex Optimization for Commitment (ROSCOC) algorithm, which directly converts its reserved exploration history into a commitment portfolio and achieves tight regret bound. Finally, we also conduct numerical experiments which confirm that our proposed algorithms achieve the desired regret predicted by our theory, and also outperform other baseline algorithms.

📄 PDF Abstract BibTeX arXiv:2607.23432

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Short-Term and Long-Term Context Aggregation Network for Video Inpainting

2020-09-12 · ECCV 2020 8 · Ang Li, Shanshan Zhao, Xingjun Ma, Mingming Gong 외

Video inpainting aims to restore missing regions of a video and has many applications such as video editing and object removal. However, existing methods either suffer from inaccurate short-term context aggregation or ra…

Video EditingVideo Inpainting

Synthesizing Long-Term Human Motions with Diffusion Models via Coherent Sampling

2023-08-03 · Zhao Yang, Bing Su, Ji-Rong Wen

Text-to-motion generation has gained increasing attention, but most existing methods are limited to generating short-term motions that correspond to a single sentence describing a single action. However, when a text stre…

Motion GenerationSentence

Efficient Pain Recognition via Respiration Signals: A Single Cross-Attention Transformer Multi-Window Fusion Pipeline

2025-07-29 · Stefanos Gkikas, Ioannis Kyprakis, Manolis Tsiknakis arxiv

Pain is a complex condition that affects a large portion of the population. Accurate and consistent evaluation is essential for individuals experiencing pain and supports the development of effective and advanced managem…

HL-OutPaint: Coarse-to-Fine Video Outpainting for High-Resolution Long-Range Videos

2026-05-17 · Jeongeun Park, Janghyeok Han, Geonung Kim, Hyun-Seung Lee 외 arxiv

Video outpainting generates plausible visual content beyond the original spatial extent of a video, playing a key role in adapting videos to diverse display formats. To support such use cases, it must enable large spatia…

Improving Pain Classification using Spatio-Temporal Deep Learning Approaches with Facial Expressions

2025-01-12 · Aafaf Ridouan, Amine Bohi, Youssef Mourchid

Pain management and severity detection are crucial for effective treatment, yet traditional self-reporting methods are subjective and may be unsuitable for non-verbal individuals (people with limited speaking skills). To…

Management