Dissecting Dissonance: Benchmarking Large Multimodal Models Against Self-Contradictory Instructions
Large multimodal models (LMMs) excel in adhering to human instructions. However, self-contradictory instructions may arise due to the increasing trend of multimodal interaction and context length, which is challenging for language beginners and vulnerable populations. We introduce the Self-Contradictory Instructions benchmark to evaluate the capability of LMMs in recognizing conflicting commands. It comprises 20,000 conflicts, evenly distributed between language and vision paradigms. It is constructed by a novel automatic dataset creation framework, which expedites the process and enables us to encompass a wide range of instruction forms. Our comprehensive evaluation reveals current LMMs consistently struggle to identify multimodal instruction discordance due to a lack of self-awareness. Hence, we propose the Cognitive Awakening Prompting to inject cognition from external, largely enhancing dissonance detection. The dataset and code are here: https://selfcontradiction.github.io/.
Code (1)
Tasks
Benchmarkingmultimodal interactionSimilar Papers 제목 키워드 기반
Topic-independent Detection of Dissonance in Short Stance Text
We address dissonance detection, the task of detecting conflicting stance between two input statements. Computational models for stance detection have typically been trained for a given target topic (e.g. gun control). I…
Stance DetectionEmpirical Evaluation of Topic Zero- and Few-Shot Learning for Stance Dissonance Detection
We address stance dissonance detection, the task of detecting conflicting stance between two input statements. Computational models for traditional stance detection have typically been trained to indicate pro/cons for a …
Few-Shot LearningStance DetectionCharacterizing Social Imaginaries and Self-Disclosures of Dissonance in Online Conspiracy Discussion Communities
Online discussion platforms offer a forum to strengthen and propagate belief in misinformed conspiracy theories. Yet, they also offer avenues for conspiracy theorists to express their doubts and experiences of cognitive …
2kBenchmarking Large Multimodal Models against Common Corruptions
This technical report aims to fill a deficiency in the assessment of large multimodal models (LMMs) by specifically examining the self-consistency of their outputs when subjected to common corruptions. We investigate the…
BenchmarkingImage to textSpeech-to-Texttext-to-speech+1From Dissonance to Insights: Dissecting Disagreements in Rationale Construction for Case Outcome Classification
In legal NLP, Case Outcome Classification (COC) must not only be accurate but also trustworthy and explainable. Existing work in explainable COC has been limited to annotations by a single expert. However, it is well-kno…