paper-with-me

Papers

Unleashing the Power of Motion and Depth: A Selective Fusion Strategy for RGB-D Video Salient Object Detection

2025-07-29 · Jiahao He, Daerji Suolang, Keren Fu, Qijun Zhao arxiv

Applying salient object detection (SOD) to RGB-D videos is an emerging task called RGB-D VSOD and has recently gained increasing interest, due to considerable performance gains of incorporating motion and depth and that RGB-D videos can be easily captured now in daily life. Existing RGB-D VSOD models have different attempts to derive motion cues, in which extracting motion information explicitly from optical flow appears to be a more effective and promising alternative. Despite this, there remains a key issue that how to effectively utilize optical flow and depth to assist the RGB modality in SOD. Previous methods always treat optical flow and depth equally with respect to model designs, without explicitly considering their unequal contributions in individual scenarios, limiting the potential of motion and depth. To address this issue and unleash the power of motion and depth, we propose a novel selective cross-modal fusion framework (SMFNet) for RGB-D VSOD, incorporating a pixel-level selective fusion strategy (PSF) that achieves optimal fusion of optical flow and depth based on their actual contributions. Besides, we propose a multi-dimensional selective attention module (MSAM) to integrate the fused features derived from PSF with the remaining RGB modality at multiple dimensions, effectively enhancing feature representation to generate refined features. We conduct comprehensive evaluation of SMFNet against 19 state-of-the-art models on both RDVS and DVisal datasets, making the evaluation the most comprehensive RGB-D VSOD benchmark up to date, and it also demonstrates the superiority of SMFNet over other models. Meanwhile, evaluation on five video benchmark datasets incorporating synthetic depth validates the efficacy of SMFNet as well. Our code and benchmark results are made publicly available at https://github.com/Jia-hao999/SMFNet.

📄 PDF Abstract BibTeX arXiv:2507.21857

Code (0)

등록된 구현이 없습니다.

Tasks

Video Salient Object Detection

Similar Papers 제목 키워드 기반

Prompting Depth Anything for 4K Resolution Accurate Metric Depth Estimation

2024-12-18 · CVPR 2025 1 · Haotong Lin, Sida Peng, Jingxiao Chen, Songyou Peng 외

Prompts play a critical role in unleashing the power of language and vision foundation models for specific tasks. For the first time, we introduce prompting into depth foundation models, creating a new paradigm for metri…

3D Reconstruction4kDecoderDepth Estimation+1

Unleashing HyDRa: Hybrid Fusion, Depth Consistency and Radar for Unified 3D Perception

2024-03-12 · Philipp Wolters, Johannes Gilg, Torben Teepe, Fabian Herzog 외

Low-cost, vision-centric 3D perception systems for autonomous driving have made significant progress in recent years, narrowing the gap to expensive LiDAR-based methods. The primary challenge in becoming a fully reliable…

3D Multi-Object Tracking3D Object Detection3D Object Detection (RoI)3D Semantic Occupancy Prediction+3

Unleashing Uncertainty: Efficient Machine Unlearning for Generative AI

2025-08-28 · Christoforos N. Spartalis, Theodoros Semertzidis, Petros Daras, Efstratios Gavves arxiv

We introduce SAFEMax, a novel method for Machine Unlearning in diffusion models. Grounded in information-theoretic principles, SAFEMax maximizes the entropy in generated images, causing the model to generate Gaussian noi…

UDPNet: Unleashing Depth-based Priors for Robust Image Dehazing

2026-01-11 · Zengyuan Zuo, Junjun Jiang, Gang Wu, Xianming Liu arxiv

Image dehazing has witnessed significant advancements with the development of deep learning models. However, most existing methods focus solely on single-modal RGB features, neglecting the inherent correlation between sc…

Computational EfficiencyDepth EstimationImage Dehazing

SELF-CARE: Selective Fusion with Context-Aware Low-Power Edge Computing for Stress Detection

2022-05-08 · Nafiul Rashid, Trier Mortlock, Mohammad Abdullah Al Faruque

Detecting human stress levels and emotional states with physiological body-worn sensors is a complex task, but one with many health-related benefits. Robustness to sensor measurement noise and energy efficiency of low-po…

Edge-computingSensor Fusion