paper-with-me

홈 › Papers

Learning When to Act: Interval-Aware Reinforcement Learning with Predictive Temporal Structure

2026-03-23 · Davide Di Gioia arxiv

Autonomous agents operating in continuous environments must decide not only what to do, but when to act. We introduce a lightweight adaptive temporal control system that learns the optimal interval between cognitive ticks from experience, replacing ad hoc biologically inspired timers with a principled learned policy. The policy state is augmented with a predictive hyperbolic spread signal (a "curvature signal" shorthand) derived from hyperbolic geometry: the mean pairwise Poincare distance among n sampled futures embedded in the Poincare ball. High spread indicates a branching, uncertain future and drives the agent to act sooner; low spread signals predictability and permits longer rest intervals. We further propose an interval-aware reward that explicitly penalises inefficiency relative to the chosen wait time, correcting a systematic credit-assignment failure of naive outcome-based rewards in timing problems. We additionally introduce a joint spatio-temporal embedding (ATCPG-ST) that concatenates independently normalised state and position projections in the Poincare ball; spatial trajectory divergence provides an independent timing signal unavailable to the state-only variant (ATCPG-SO). This extension raises mean hyperbolic spread (kappa) from 1.88 to 3.37 and yields a further 5.8 percent efficiency gain over the state-only baseline. Ablation experiments across five random seeds demonstrate that (i) learning is the dominant efficiency factor (54.8 percent over no-learning), (ii) hyperbolic spread provides significant complementary gain (26.2 percent over geometry-free control), (iii) the combined system achieves 22.8 percent efficiency over the fixed-interval baseline, and (iv) adding spatial position information to the spread embedding yields an additional 5.8 percent.

📄 PDF Abstract BibTeX arXiv:2603.22384

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning

Similar Papers 제목 키워드 기반

Time-Aware Q-Networks: Resolving Temporal Irregularity for Deep Reinforcement Learning

2021-05-06 · Yeo Jin Kim, Min Chi

Deep Reinforcement Learning (DRL) has shown outstanding performance on inducing effective action policies that maximize expected long-term return on many complex tasks. Much of DRL work has been focused on sequences of e…

Deep Reinforcement Learningreinforcement-learningReinforcement Learning (RL)Temporal Sequences

TSCoNet: A Two-Stage Copula CNN-LSTM for Uncertainty-Aware Spatio-Temporal Forecasting

2026-07-11 · Jongwook Kim, Jong-Min Kim arxiv

Reliable forecasting of several interrelated environmental variables - such as regional precipitation and temperature, or other correlated geophysical fields - across many locations calls for accurate predictions accompa…

Auto-AEG: Scalable Data Construction for Open-Vocabulary Audio Event Grounding

2026-07-05 · Zihan Zhang, Xize Cheng, Wenhao Yan, Tong Zhang 외 arxiv

Large Audio-Language Models (LALMs) reason fluently about sound yet struggle to localize precisely when events occur, while classical Sound Event Detection attains frame-level precision only over a closed label set. At t…

Reinforcement LearningSound Event Detection

Temporal Inversion for Learning Interval Change in Chest X-Rays

2026-04-06 · Hanbin Ko, Kyungmin Jeon, Doowoong Choi, Chang Min Park arxiv

Recent advances in vision--language pretraining have enabled strong medical foundation models, yet most analyze radiographs in isolation, overlooking the key clinical task of comparing prior and current images to assess …

TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning

2025-12-03 · Tao Wu, Li Yang, Gen Zhan, Yabin Zhang 외 arxiv

Enhancing the temporal understanding of Multimodal Large Language Models (MLLMs) is essential for advancing long-form video analysis, enabling tasks such as temporal localization, action detection, and time-sensitive que…

Reinforcement LearningQuestion AnsweringAction Detection