paper-with-me

홈 › Papers

C^3Net: End-to-End deep learning for efficient real-time visual active camera control

2021-07-28 · Christos Kyrkou

The need for automated real-time visual systems in applications such as smart camera surveillance, smart environments, and drones necessitates the improvement of methods for visual active monitoring and control. Traditionally, the active monitoring task has been handled through a pipeline of modules such as detection, filtering, and control. However, such methods are difficult to jointly optimize and tune their various parameters for real-time processing in resource constraint systems. In this paper a deep Convolutional Camera Controller Neural Network is proposed to go directly from visual information to camera movement to provide an efficient solution to the active vision problem. It is trained end-to-end without bounding box annotations to control a camera and follow multiple targets from raw pixel values. Evaluation through both a simulation framework and real experimental setup, indicate that the proposed solution is robust to varying conditions and able to achieve better monitoring performance than traditional approaches both in terms of number of targets monitored as well as in effective monitoring time. The advantage of the proposed approach is that it is computationally less demanding and can run at over 10 FPS (~4x speedup) on an embedded smart camera providing a practical and affordable solution to real-time active monitoring.

📄 PDF Abstract BibTeX arXiv:2107.13233

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Imitation-Based Active Camera Control with Deep Convolutional Neural Network

2020-12-11 · Christos Kyrkou

The increasing need for automated visual monitoring and control for applications such as smart camera surveillance, traffic monitoring, and intelligent environments, necessitates the improvement of methods for visual act…

Imitation Learning

4DStreamCtrl: Interactive Video Generation with Online 4D Control

2026-08-26 · Shiqian Li, Chenguo Lin, Zhiguang Liu, Yu Tang 외 arxiv

Generative video models now synthesize footage nearly indistinguishable from reality. Their promise as interactive tools hinges on fine-grained control of how objects and the camera move over time, yet each existing appr…

Video GenerationCausal Inference

RealCam: Real-Time Novel-View Video Generation with Interactive Camera Control

2026-05-07 · Youcan Xu, Jiaxin Shi, Zhen Wang, Wensong Song 외 arxiv

Camera-controlled video-to-video (V2V) generation enables dynamic viewpoint synthesis from monocular footage, holding immense potential for interactive filmmaking and live broadcasting. However, existing implicit synthes…

Data AugmentationVideo Generation

Look, Zoom, Understand: The Robotic Eyeball for Embodied Perception

2025-11-19 · Jiashu Yang, Yifan Han, Yucheng Xie, Ning Guo 외 arxiv

In embodied AI, visual perception should be active rather than passive: the system must decide where to look and at what scale to sense to acquire maximally informative data under pixel and spatial budget constraints. Ex…

Reinforcement LearningMultimodal Reasoning

Towards Active Vision for Action Localization with Reactive Control and Predictive Learning

2021-11-09 · Shubham Trehan, Sathyanarayanan N. Aakur

Visual event perception tasks such as action localization have primarily focused on supervised learning settings under a static observer, i.e., the camera is static and cannot be controlled by an algorithm. They are ofte…

Action LocalizationDiversityObject Tracking