paper-with-me

홈 › Papers

Addressing Missing and Noisy Modalities in One Solution: Unified Modality-Quality Framework for Low-quality Multimodal Data

2026-03-03 · Sijie Mai, Shiqin Han, Haifeng Hu arxiv

Multimodal data encountered in real-world scenarios are typically of low quality, with noisy modalities and missing modalities being typical forms that severely hinder model performance and robustness. However, prior works often handle noisy and missing modalities separately. In contrast, we jointly address missing and noisy modalities to enhance model robustness in low-quality data scenarios. We regard both noisy and missing modalities as a unified low-quality modality problem, and propose a unified modality-quality (UMQ) framework to enhance low-quality representations for multimodal affective computing. Firstly, we train a quality estimator with explicit supervised signals via a rank-guided training strategy that compares the relative quality of different representations by adding a ranking constraint, avoiding training noise caused by inaccurate absolute quality labels. Then, a quality enhancer for each modality is constructed, which uses the sample-specific information provided by other modalities and the modality-specific information provided by the defined modality baseline representation to enhance the quality of unimodal representations. Finally, we propose a quality-aware mixture-of-experts module with particular routing mechanism to enable multiple modality-quality problems to be addressed more specifically. UMQ consistently outperforms state-of-the-art baselines on multiple datasets under the settings of complete, missing, and noisy modalities.

📄 PDF Abstract BibTeX arXiv:2603.02695

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

DCER: Dual-Stage Compression and Energy-Based Reconstruction

2026-02-03 · Yiwen Wang, Jiahao Qin arxiv

Multimodal fusion faces two robustness challenges: noisy inputs degrade representation quality, and missing modalities cause prediction failures. We propose DCER, a unified framework addressing both challenges through du…

Multimodal Sleep Apnea Detection with Missing or Noisy Modalities

2024-02-24 · Hamed Fayyaz, Abigail Strang, Niharika S. D'Souza, Rahmatollah Beheshti

Polysomnography (PSG) is a type of sleep study that records multimodal physiological signals and is widely used for purposes such as sleep staging and respiratory event detection. Conventional machine learning methods as…

Event DetectionSleep apnea detectionSleep Staging

Towards Stable Cross-Domain Depression Recognition under Missing Modalities

2025-12-06 · Jiuyi Chen, Mingkui Tan, Haifeng Lu, Qiuna Xu 외 arxiv

Depression poses serious public health risks, including suicide, underscoring the urgency of timely and scalable screening. Multimodal automatic depression detection (ADD) offers a promising solution; however, widely stu…

Domain Generalization

TTSD-FAR: Test-Time Self-Distillation with Fisher-Anchored Restoration for Missing-Modality Emotion Recognition in LVLMs

2026-08-18 · Muhammad Haseeb Aslam, Alessandro Koerich, Marco Pedersoli, Ali Etemad 외 arxiv

Large video-language models (LVLMs) have shown remarkable performance on multimodal tasks like multimodal emotion recognition (ER) in the wild. ER is inherently multimodal, requiring a joint understanding of facial expre…

Multimodal Emotion Recognition

Speech Translation with Speech Foundation Models and Large Language Models: What is There and What is Missing?

2024-02-19 · Marco Gaido, Sara Papi, Matteo Negri, Luisa Bentivogli

The field of natural language processing (NLP) has recently witnessed a transformative shift with the emergence of foundation models, particularly Large Language Models (LLMs) that have revolutionized text-based NLP. Thi…

Speech-to-TextSpeech-to-Text Translation