paper-with-me

Papers

Model-guided Multi-path Knowledge Aggregation for Aerial Saliency Prediction

2018-11-14 · Kui Fu, Jia Li, Yu Zhang, Hongze Shen, Yonghong Tian

As an emerging vision platform, a drone can look from many abnormal viewpoints which brings many new challenges into the classic vision task of video saliency prediction. To investigate these challenges, this paper proposes a large-scale video dataset for aerial saliency prediction, which consists of ground-truth salient object regions of 1,000 aerial videos, annotated by 24 subjects. To the best of our knowledge, it is the first large-scale video dataset that focuses on visual saliency prediction on drones. Based on this dataset, we propose a Model-guided Multi-path Network (MM-Net) that serves as a baseline model for aerial video saliency prediction. Inspired by the annotation process in eye-tracking experiments, MM-Net adopts multiple information paths, each of which is initialized under the guidance of a classic saliency model. After that, the visual saliency knowledge encoded in the most representative paths is selected and aggregated to improve the capability of MM-Net in predicting spatial saliency in aerial scenarios. Finally, these spatial predictions are adaptively combined with the temporal saliency predictions via a spatiotemporal optimization algorithm. Experimental results show that MM-Net outperforms ten state-of-the-art models in predicting aerial video saliency.

📄 PDF Abstract BibTeX arXiv:1811.05625

Code (0)

등록된 구현이 없습니다.

Tasks

Aerial Video Saliency PredictionPredictionSaliency PredictionTransfer LearningVideo Saliency Prediction

Similar Papers 제목 키워드 기반

CrossVL: Complexity-Aware Feature Routing and Paired Curriculum for Cross-View Vision-Language Detection

2026-05-10 · Zhipeng Liu, Chunbo Luo arxiv

Vision-language models (VLMs) enable text-guided object detection but degrade severely under cross-view scenarios where ground and aerial viewpoints differ in altitude, scale, and spatial layout. These geometric changes …

Object Detection

Aerial Multi-View Stereo via Adaptive Depth Range Inference and Normal Cues

2025-06-06 · Yimei Liu, Yakun Ju, Yuan Rao, Hao Fan 외

Three-dimensional digital urban reconstruction from multi-view aerial images is a critical application where deep multi-view stereo (MVS) methods outperform traditional techniques. However, existing methods commonly over…

Depth Estimation

SliceMatch: Geometry-guided Aggregation for Cross-View Pose Estimation

2022-11-26 · CVPR 2023 1 · Ted Lentsch, Zimin Xia, Holger Caesar, Julian F. P. Kooij

This work addresses cross-view camera pose estimation, i.e., determining the 3-Degrees-of-Freedom camera pose of a given ground-level image w.r.t. an aerial image of the local area. We propose SliceMatch, which consists …

Camera Pose EstimationContrastive Learningfeature selectionPose Estimation+1

SkyRover: A Modular Simulator for Cross-Domain Pathfinding

2025-02-13 · Wenhui Ma, Wenhao Li, Bo Jin, Changhong Lu 외

Unmanned Aerial Vehicles (UAVs) and Automated Guided Vehicles (AGVs) increasingly collaborate in logistics, surveillance, inspection tasks and etc. However, existing simulators often focus on a single domain, limiting cr…

Benchmarking

LeanRAG: Knowledge-Graph-Based Generation with Semantic Aggregation and Hierarchical Retrieval

2025-08-14 · Yaoze Zhang, Rong Wu, Pinlong Cai, Xiaoman Wang 외 arxiv

Retrieval-Augmented Generation (RAG) plays a crucial role in grounding Large Language Models by leveraging external knowledge, whereas the effectiveness is often compromised by the retrieval of contextually flawed or inc…

Information Retrieval