paper-with-me

홈 › Papers

General Incomplete Multimodal Learning via Dynamic Quality Perception

2026-07-08 · Xiangyu Meng, Shicai Wei arxiv

Multimodal learning robust to missing modalities is essential for real-world applications. Existing methods mainly focus on inter-modality missing, where entire modalities are absent, while overlooking intra-modality degradation, where modalities are present but severely corrupted. In practice, these two types of missing often coexist, making existing approaches ineffective. To address this limitation, we propose General Incomplete Multimodal Learning (GIML), a unified framework that simultaneously handles both inter-modality missing and intra-modality degradation through dynamic quality perception. Specifically, GIML models heterogeneous missing patterns as continuous modality information degradation, enabling degradation-aware adaptive fusion. To achieve reliable quality perception, we introduce a Noise-aware Quality Estimator that learns the mapping from corrupted features to noise intensity through controlled noise injection. Furthermore, we propose a Noise-Semantic Decoupled module that separates semantic information from noise interference. This improves robustness and generalization to unseen corruption patterns. Extensive experiments across datasets with diverse modality types demonstrate the effectiveness and generality of GIML. Code is available at: https://github.com/Yu-Five/GIML.

📄 PDF Abstract BibTeX arXiv:2607.06943

Code (1)

Tavish9/awesome-daily-AI-arxiv ★ 111

Similar Papers 제목 키워드 기반

VideoAesBench: Benchmarking the Video Aesthetics Perception Capabilities of Large Multimodal Models

2026-01-29 · Yunhao Li, Sijing Wu, Zhilin Gao, Zicheng Zhang 외 arxiv

Large multimodal models (LMMs) have demonstrated outstanding capabilities in various visual perception tasks, which has in turn made the evaluation of LMMs significant. However, the capability of video aesthetic quality …

LIMSSR: LLM-Driven Sequence-to-Score Reasoning under Training-Time Incomplete Multimodal Observations

2026-05-01 · Huangbiao Xu, Huanqi Wu, Xiao Ke, Yuxin Peng arxiv

Real-world multimodal learning is often hindered by missing modalities. While Incomplete Multimodal Learning (IML) has gained traction, existing methods typically rely on the unrealistic assumption of full-modal availabi…

Action Quality Assessment

Depth-Semantic Alignment and Affinity-Guided Fusion for Structured Radar Point Cloud Generation

2026-06-25 · Amjad Hussain, Xin Qiu, Fuyuan Ai, Yuchen Tan 외 arxiv

Point clouds are an important carrier of three-dimensional spatial information, and their quality directly affects the performance of downstream perception tasks such as object detection and tracking. However, millimeter…

Point Cloud GenerationObject DetectionPoint Clouds

City-VLM: Towards Multidomain Perception Scene Understanding via Multimodal Incomplete Learning

2025-07-17 · Penglei Sun, Yaoxian Song, Xiangru Zhu, Xiang Liu 외

Scene understanding enables intelligent agents to interpret and comprehend their environment. While existing large vision-language models (LVLMs) for scene understanding have primarily focused on indoor household tasks, …

Question AnsweringScene Understanding

Perceptual 3D Simulation With Physical World Modeling

2026-06-25 · Wanhee Lee, Klemen Kotar, Rahul Mysore Venkatesh, Jared Watrous 외 arxiv

Predicting how a scene will evolve after a desired 3D transformation from images is a central goal in vision, graphics, and robotics. Yet unlike ideal simulators with full access to 3D geometry and dynamics, real world s…

Novel View SynthesisScene Understanding