paper-with-me

홈 › Papers

Detecting Multimodal Situations with Insufficient Context and Abstaining from Baseless Predictions

2024-05-18 · Junzhang Liu, Zhecan Wang, Hammad Ayyubi, Haoxuan You, Chris Thomas, Rui Sun, Shih-Fu Chang, Kai-Wei Chang

Despite the widespread adoption of Vision-Language Understanding (VLU) benchmarks such as VQA v2, OKVQA, A-OKVQA, GQA, VCR, SWAG, and VisualCOMET, our analysis reveals a pervasive issue affecting their integrity: these benchmarks contain samples where answers rely on assumptions unsupported by the provided context. Training models on such data foster biased learning and hallucinations as models tend to make similar unwarranted assumptions. To address this issue, we collect contextual data for each sample whenever available and train a context selection module to facilitate evidence-based model predictions. Strong improvements across multiple benchmarks demonstrate the effectiveness of our approach. Further, we develop a general-purpose Context-AwaRe Abstention (CARA) detector to identify samples lacking sufficient context and enhance model accuracy by abstaining from responding if the required context is absent. CARA exhibits generalization to new benchmarks it wasn't trained on, underscoring its utility for future VLU benchmarks in detecting or cleaning samples with inadequate context. Finally, we curate a Context Ambiguity and Sufficiency Evaluation (CASE) set to benchmark the performance of insufficient context detectors. Overall, our work represents a significant advancement in ensuring that vision-language models generate trustworthy and evidence-based outputs in complex real-world scenarios.

📄 PDF Abstract BibTeX arXiv:2405.11145

Code (0)

등록된 구현이 없습니다.

Tasks

Visual Question Answering (VQA)

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically

Similar Papers 제목 키워드 기반

SmartRSD: An Intelligent Multimodal Approach to Real-Time Road Surface Detection for Safe Driving

2024-06-14 · Adnan Md Tayeb, Mst Ayesha Khatun, Mohtasin Golam, Md Facklasur Rahaman 외

Precise and prompt identification of road surface conditions enables vehicles to adjust their actions, like changing speed or using specific traction control techniques, to lower the chance of accidents and potential dan…

Global explainability of a deep abstaining classifier

2025-04-01 · Sayera Dhaubhadel, Jamaludin Mohd-Yusof, Benjamin H. McMahon, Trilce Estrada 외

We present a global explainability method to characterize sources of errors in the histology prediction task of our real-world multitask convolutional neural network (MTCNN)-based deep abstaining classifier (DAC), for au…

Dimensionality Reduction

Multimodal Contextual Dialogue Breakdown Detection for Conversational AI Models

2024-04-11 · Md Messal Monem Miah, Ulie Schnaithmann, Arushi Raghuvanshi, Youngseo Son

Detecting dialogue breakdown in real time is critical for conversational AI systems, because it enables taking corrective action to successfully complete a task. In spoken dialog systems, this breakdown can be caused by …

Navigate

Uncertainty-Based Abstention in LLMs Improves Safety and Reduces Hallucinations

2024-04-16 · Christian Tomani, Kamalika Chaudhuri, Ivan Evtimov, Daniel Cremers 외

A major barrier towards the practical deployment of large language models (LLMs) is their lack of reliability. Three situations where this is particularly apparent are correctness, hallucinations when given unanswerable …

Question Answering

Neuromorphic Drone Detection: an Event-RGB Multimodal Approach

2024-09-24 · Gabriele Magrini, Federico Becattini, Pietro Pala, Alberto del Bimbo 외

In recent years, drone detection has quickly become a subject of extreme interest: the potential for fast-moving objects of contained dimensions to be used for malicious intents or even terrorist attacks has posed attent…

object-detectionObject Detection