paper-with-me

Papers

Calibrating Multimodal Learning

2023-06-02 · Huan Ma. Qingyang Zhang, Changqing Zhang, Bingzhe Wu, Huazhu Fu, Joey Tianyi Zhou, QinGhua Hu

Multimodal machine learning has achieved remarkable progress in a wide range of scenarios. However, the reliability of multimodal learning remains largely unexplored. In this paper, through extensive empirical studies, we identify current multimodal classification methods suffer from unreliable predictive confidence that tend to rely on partial modalities when estimating confidence. Specifically, we find that the confidence estimated by current models could even increase when some modalities are corrupted. To address the issue, we introduce an intuitive principle for multimodal learning, i.e., the confidence should not increase when one modality is removed. Accordingly, we propose a novel regularization technique, i.e., Calibrating Multimodal Learning (CML) regularization, to calibrate the predictive confidence of previous methods. This technique could be flexibly equipped by existing models and improve the performance in terms of confidence calibration, classification accuracy, and model robustness.

📄 PDF Abstract BibTeX arXiv:2306.01265

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Proceedings of the 40th International Conference on Machine Learning

2023-07-01 · journal 2023 7 · Huan Ma, Qingyang Zhang, Changqing Zhang, Bingzhe Wu 외

Multimodal machine learning has achieved remarkable progress in a wide range of scenarios. However, the reliability of multimodal learning remains largely unexplored. In this paper, through extensive empirical studies, w…

IMG: Calibrating Diffusion Models via Implicit Multimodal Guidance

2025-09-30 · Jiayi Guo, Chuanhao Yan, Xingqian Xu, Yulin Wang 외 arxiv

Ensuring precise multimodal alignment between diffusion-generated images and input prompts has been a long-standing challenge. Earlier works finetune diffusion weight using high-quality preference data, which tends to be…

Calibrating Multimodal Consensus for Emotion Recognition

2025-10-23 · Guowei Zhong, Junjie Li, Huaiyu Zhu, Ruohong Huan 외 arxiv

In recent years, Multimodal Emotion Recognition (MER) has made substantial progress. Nevertheless, most existing approaches neglect the semantic inconsistencies that may arise across modalities, such as conflicting emoti…

Multimodal Emotion Recognition

Calibrating MLLM-as-a-judge via Multimodal Bayesian Prompt Ensembles

2025-09-10 · Eric Slyman, Mehrab Tanjim, Kushal Kafle, Stefan Lee arxiv

Multimodal large language models (MLLMs) are increasingly used to evaluate text-to-image (TTI) generation systems, providing automated judgments based on visual and textual context. However, these "judge" models often su…

Image Clustering

Balanced Multimodal Learning via Mutual Information

2025-11-02 · Rongrong Xie, Guido Sanguinetti arxiv

Multimodal learning has increasingly become a focal point in research, primarily due to its ability to integrate complementary information from diverse modalities. Nevertheless, modality imbalance, stemming from factors …

Knowledge Distillation