paper-with-me

홈 › Papers

Adaptive Confidence Regularization for Multimodal Failure Detection

2026-03-02 · Moru Liu, Hao Dong, Olga Fink, Mario Trapp arxiv

The deployment of multimodal models in high-stakes domains, such as self-driving vehicles and medical diagnostics, demands not only strong predictive performance but also reliable mechanisms for detecting failures. In this work, we address the largely unexplored problem of failure detection in multimodal contexts. We propose Adaptive Confidence Regularization (ACR), a novel framework specifically designed to detect multimodal failures. Our approach is driven by a key observation: in most failure cases, the confidence of the multimodal prediction is significantly lower than that of at least one unimodal branch, a phenomenon we term confidence degradation. To mitigate this, we introduce an Adaptive Confidence Loss that penalizes such degradations during training. In addition, we propose Multimodal Feature Swapping, a novel outlier synthesis technique that generates challenging, failure-aware training examples. By training with these synthetic failures, ACR learns to more effectively recognize and reject uncertain predictions, thereby improving overall reliability. Extensive experiments across four datasets, three modalities, and multiple evaluation settings demonstrate that ACR achieves consistent and robust gains. The source code will be available at https://github.com/mona4399/ACR.

📄 PDF Abstract BibTeX arXiv:2603.02200

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Calibrating Multimodal Learning

2023-06-02 · Huan Ma. Qingyang Zhang, Changqing Zhang, Bingzhe Wu, Huazhu Fu 외

Multimodal machine learning has achieved remarkable progress in a wide range of scenarios. However, the reliability of multimodal learning remains largely unexplored. In this paper, through extensive empirical studies, w…

Proceedings of the 40th International Conference on Machine Learning

2023-07-01 · journal 2023 7 · Huan Ma, Qingyang Zhang, Changqing Zhang, Bingzhe Wu 외

Multimodal machine learning has achieved remarkable progress in a wide range of scenarios. However, the reliability of multimodal learning remains largely unexplored. In this paper, through extensive empirical studies, w…

Evaluating Uncertainty-based Failure Detection for Closed-Loop LLM Planners

2024-06-01 · Zhi Zheng, Qian Feng, Hang Li, Alois Knoll 외

Recently, Large Language Models (LLMs) have witnessed remarkable performance as zero-shot task planners for robotic manipulation tasks. However, the open-loop nature of previous works makes LLM-based planning error-prone…

RFM-HRI : A Multimodal Dataset of Medical Robot Failure, User Reaction and Recovery Preferences for Item Retrieval Tasks

2026-03-05 · Yashika Batra, Giuliano Pioldi, Promise Ekpo, Arman Sayatqyzy 외 arxiv

While robots deployed in real-world environments inevitably experience interaction failures, understanding how users respond through verbal and non-verbal behaviors remains under-explored in human-robot interaction (HRI)…

Dynamic Gated Cross-Modal Fusion with Sarcastic-aware Contrastive Regularization for Multimodal Sarcasm Detection

2026-08-20 · Hao Guo, Subin Huang, Junjie Chen, Zhifa Geng 외 arxiv

Multimodal sarcasm detection aims to identify sarcastic intent from multimodal content, where inconsistencies between literal meaning and contextual cues often signal irony. This task has attracted increasing research at…

Sarcasm Detection