paper-with-me

Papers

DiffVAS: Diffusion-Guided Visual Active Search in Partially Observable Environments

2026-05-15 · Anindya Sarkar, Srikumar Sastry, Aleksis Pirinen, Nathan Jacobs, Yevgeniy Vorobeychik arxiv

Visual active search (VAS) has been introduced as a modeling framework that leverages visual cues to direct aerial (e.g., UAV-based) exploration and pinpoint areas of interest within extensive geospatial regions. Potential applications of VAS include detecting hotspots for rare wildlife poaching, aiding search-and-rescue missions, and uncovering illegal trafficking of weapons, among other uses. Previous VAS approaches assume that the entire search space is known upfront, which is often unrealistic due to constraints such as a restricted field of view and high acquisition costs, and they typically learn policies tailored to specific target objects, which limits their ability to search for multiple target categories simultaneously. In this work, we propose DiffVAS, a target-conditioned policy that searches for diverse objects simultaneously according to task requirements in partially observable environments, which advances the deployment of visual active search policies in real-world applications. DiffVAS leverages a diffusion model to reconstruct the entire geospatial area from sequentially observed partial glimpses, which enables a target-conditioned reinforcement learning-based planning module to effectively reason and guide subsequent search steps. Extensive experiments demonstrate that DiffVAS excels in searching diverse objects in partially observable environments, significantly surpassing state-of-the-art methods on several datasets.

📄 PDF Abstract BibTeX arXiv:2605.15519

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning

Similar Papers 제목 키워드 기반

SurGen: Text-Guided Diffusion Model for Surgical Video Generation

2024-08-26 · Joseph Cho, Samuel Schmidgall, Cyril Zakka, Mrudang Mathur 외

Diffusion-based video generation models have made significant strides, producing outputs with improved visual fidelity, temporal coherence, and user control. These advancements hold great promise for improving surgical e…

Video Generation

ARDuP: Active Region Video Diffusion for Universal Policies

2024-06-19 · Shuaiyi Huang, Mara Levy, Zhenyu Jiang, Anima Anandkumar 외

Sequential decision-making can be formulated as a text-conditioned video generation problem, where a video planner, guided by a text-defined goal, generates future frames visualizing planned actions, from which control a…

Decision MakingSequential Decision MakingVideo Generation

SwiftPie: Lightning-fast Subject-driven Image Personalization via One step Diffusion

2026-05-02 · Huy Duong, Trong-Tung Nguyen, Cuong Pham, Anh Tran 외 arxiv

Diffusion models have achieved remarkable success in high-quality image synthesis, sparking interest in image-guided generation tasks such as subject-driven image personalization. Despite their impressive personalization…

Personalized Image Generation

DiT as Real-Time Rerenderer: Streaming Video Stylization with Autoregressive Diffusion Transformer

2026-04-15 · Hengye Lyu, Zisu Li, Yue Hong, Yueting Weng 외 arxiv

Recent advances in video generation models has significantly accelerated video generation and related downstream tasks. Among these, video stylization holds important research value in areas such as immersive application…

Video Generation

DDP: Diffusion Model for Dense Visual Prediction

2023-03-30 · ICCV 2023 1 · Yuanfeng Ji, Zhe Chen, Enze Xie, Lanqing Hong 외

We propose a simple, efficient, yet powerful framework for dense visual predictions based on the conditional diffusion pipeline. Our approach follows a "noise-to-map" generative paradigm for prediction by progressively r…

DenoisingDepth EstimationmodelMonocular Depth Estimation+3