AnyMod-LLVE: Low-Light Video Enhancement with Modality-Agnostic Inference
Low-light video enhancement (LLVE) remains a challenging task due to severe information degradation under low-illumination conditions. Recent multimodal approaches have significantly improved enhancement performance by incorporating auxiliary modalities, such as event streams and infrared images. However, these methods typically assume the availability of these modalities at inference, which is often not feasible in real-world scenarios. To solve this problem, in this work, we propose AMNet, a unified multimodal framework for LLVE, to support flexible modality-agnostic inference, where auxiliary modalities may be unavailable. To address the issue of modality absence, we introduce a Spatial-Spectral Dual-Gated Translator that learns the correspondence between auxiliary modalities and RGB inputs, producing implicit auxiliary representations to support the robust enhancement. Additionally, to fully facilitate the learning of cross-modal correspondence, we conduct large-scale multimodal pretraining based on the RGB-only dataset with synthetic auxiliary modalities. Extensive experiments demonstrate that AMNet could handle arbitrary inference-time modality combinations and exhibits superior performance for LLVE under modality absence conditions. Code and models are available on the project page.
Code (0)
등록된 구현이 없습니다.
Tasks
Video EnhancementSimilar Papers 제목 키워드 기반
Light-VQA: A Multi-Dimensional Quality Assessment Model for Low-Light Video Enhancement
Recently, Users Generated Content (UGC) videos becomes ubiquitous in our daily lives. However, due to the limitations of photographic equipments and techniques, UGC videos often contain various degradations, in which one…
Video EnhancementVideo Quality AssessmentVisual Question Answering (VQA)FastLLVE: Real-Time Low-Light Video Enhancement with Intensity-Aware Lookup Table
Low-Light Video Enhancement (LLVE) has received considerable attention in recent years. One of the critical requirements of LLVE is inter-frame brightness consistency, which is essential for maintaining the temporal cohe…
Video EnhancementLow-Light Video Enhancement with An Effective Spatial-Temporal Decomposition Paradigm
Low-Light Video Enhancement (LLVE) seeks to restore dynamic or static scenes plagued by severe invisibility and noise. In this paper, we present an innovative video decomposition strategy that incorporates view-independe…
Video EnhancementLow-Light Video Enhancement with Synthetic Event Guidance
Low-light video enhancement (LLVE) is an important yet challenging task with many applications such as photographing and autonomous driving. Unlike single image low-light enhancement, most LLVE methods utilize temporal i…
Autonomous DrivingImage EnhancementVideo EnhancementLight-VQA+: A Video Quality Assessment Model for Exposure Correction with Vision-Language Guidance
Recently, User-Generated Content (UGC) videos have gained popularity in our daily lives. However, UGC videos often suffer from poor exposure due to the limitations of photographic equipment and techniques. Therefore, Vid…
Exposure CorrectionVideo EnhancementVideo Quality AssessmentVisual Question Answering (VQA)