paper-with-me

홈 › Papers

Multi-Modal Synergistic Implicit Image Enhancement for Efficient Optical Flow Estimation

2025-01-01 · CVPR 2025 1 · Weichen Dai, Hexing Wu, Xiaoyang Weng, Yuxin Zheng, Yuhang Ming, Wanzeng Kong

As a fundamental visual task, optical flow estimation has widespread applications in computer vision. However, it faces significant challenges under adverse lighting conditions, where low texture and noise make accurate optical flow estimation particularly difficult.In this paper, we propose an optical flow method that employs implicit image enhancement through multi-modal synergistic training. To supplement the scene information missing in the original low-quality image, we utilize a high-low frequency feature enhancement network. The enhancement network is implicitly guided by multi-modal data and the specific subsequent tasks, enabling the model to learn multi-modal knowledge that enhances feature information suitable for optical flow estimation during inference. By using RGBD multi-modal data, the proposed method avoids the reliance on the images captured from the same view, a common limitation in traditional image enhancement methods.During training, the encoded features extracted from the enhanced images are synergistically supervised by features from the RGBD fusion as well as by the optical flow task.Experiments conducted on both synthetic and real datasets demonstrate that the proposed method significantly improves performance on public datasets.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Image EnhancementOptical Flow Estimation

Similar Papers 제목 키워드 기반

CtrlFuse: Mask-Prompt Guided Controllable Infrared and Visible Image Fusion

2026-01-12 · Yiming Sun, Yuan Ruan, Qinghua Hu, Pengfei Zhu arxiv

Infrared and visible image fusion generates all-weather perception-capable images by combining complementary modalities, enhancing environmental awareness for intelligent unmanned systems. Existing methods either focus o…

Synergistic Information Disentanglement for Omni-modal Slide Representation Learning in Computational Pathology

2026-09-02 · Mingxin Liu, Chengfei Cai, Anwen Lu, Pengbo Xu 외 arxiv

In computational pathology (CPath), developing omni-modal self-supervised learning (SSL) models that integrate histology, genomics, and clinical reports enables transferable representation learning for whole slide images…

Self-Supervised LearningRepresentation Learning

Connecting Giants: Synergistic Knowledge Transfer of Large Multimodal Models for Few-Shot Learning

2025-10-13 · Hao Tang, Shengfeng He, Jing Qin arxiv

Few-shot learning (FSL) addresses the challenge of classifying novel classes with limited training samples. While some methods leverage semantic knowledge from smaller-scale models to mitigate data scarcity, these approa…

Few-Shot Learning

Hire: Hybrid-modal Interaction with Multiple Relational Enhancements for Image-Text Matching

2024-06-05 · Xuri Ge, Fuhai Chen, Songpei Xu, Fuxiang Tao 외

Image-text matching (ITM) is a fundamental problem in computer vision. The key issue lies in jointly learning the visual and textual representation to estimate their similarity accurately. Most existing methods focus on …

cross-modal alignmentImage-text matchingRelationship DetectionText Matching

Dual-Domain Perspective on Degradation-Aware Fusion: A VLM-Guided Robust Infrared and Visible Image Fusion Framework

2025-09-05 · Tianpei Zhang, Jufeng Zhao, Yiming Zhu, Guangmang Cui arxiv

Most existing infrared-visible image fusion (IVIF) methods assume high-quality inputs, and therefore struggle to handle dual-source degraded scenarios, typically requiring manual selection and sequential application of m…