Towards Active Vision for Action Localization with Reactive Control and Predictive Learning
Visual event perception tasks such as action localization have primarily focused on supervised learning settings under a static observer, i.e., the camera is static and cannot be controlled by an algorithm. They are often restricted by the quality, quantity, and diversity of \textit{annotated} training data and do not often generalize to out-of-domain samples. In this work, we tackle the problem of active action localization where the goal is to localize an action while controlling the geometric and physical parameters of an active camera to keep the action in the field of view without training data. We formulate an energy-based mechanism that combines predictive learning and reactive control to perform active action localization without rewards, which can be sparse or non-existent in real-world environments. We perform extensive experiments in both simulated and real-world environments on two tasks - active object tracking and active action localization. We demonstrate that the proposed approach can generalize to different tasks and environments in a streaming fashion, without explicit rewards or training. We show that the proposed approach outperforms unsupervised baselines and obtains competitive performance compared to those trained with reinforcement learning.
Code (1)
Tasks
Action LocalizationDiversityObject TrackingSimilar Papers 제목 키워드 기반
ReactiveGWM: Steering NPC in Reactive Game World Models
Current game world models simulate environments from a subjective, player-centric perspective. However, by treating the Non-Player Character (NPC) merely as background pixels, these models cannot capture interactions bet…
Dual Policy Iteration
Recently, a novel class of Approximate Policy Iteration (API) algorithms have demonstrated impressive practical performance (e.g., ExIt from [2], AlphaGo-Zero from [27]). This new family of algorithms maintains, and alte…
continuous-controlContinuous ControlHierarchical Provision of Distribution Grid Flexibility with Online Feedback Optimization
Utilizing distribution grid flexibility for ancillary services requires the coordination and dispatch of requested active and reactive power to a large number of distributed energy resources in underlying grid layers. Th…
Computational EfficiencyPoint TrackingOptimal Selective Attention in Reactive Agents
In POMDPs, information about the hidden state, delivered through observations, is both valuable to the agent, allowing it to base its actions on better informed internal states, and a "curse", exploding the size and dive…
DiversityNeural network algorithm and its application in reactive distillation
Reactive distillation is a special distillation technology based on the coupling of chemical reaction and distillation. It has the characteristics of low energy consumption and high separation efficiency. However, becaus…