paper-with-me

홈 › Papers

iReason: Multimodal Commonsense Reasoning using Videos and Natural Language with Interpretability

2021-06-25 · Aman Chadha, Vinija Jain

Causality knowledge is vital to building robust AI systems. Deep learning models often perform poorly on tasks that require causal reasoning, which is often derived using some form of commonsense knowledge not immediately available in the input but implicitly inferred by humans. Prior work has unraveled spurious observational biases that models fall prey to in the absence of causality. While language representation models preserve contextual knowledge within learned embeddings, they do not factor in causal relationships during training. By blending causal relationships with the input features to an existing model that performs visual cognition tasks (such as scene understanding, video captioning, video question-answering, etc.), better performance can be achieved owing to the insight causal relationships bring about. Recently, several models have been proposed that have tackled the task of mining causal data from either the visual or textual modality. However, there does not exist widespread research that mines causal relationships by juxtaposing the visual and language modalities. While images offer a rich and easy-to-process resource for us to mine causality knowledge from, videos are denser and consist of naturally time-ordered events. Also, textual information offers details that could be implicit in videos. We propose iReason, a framework that infers visual-semantic commonsense knowledge using both videos and natural language captions. Furthermore, iReason's architecture integrates a causal rationalization module to aid the process of interpretability, error analysis and bias detection. We demonstrate the effectiveness of iReason using a two-pronged comparative analysis with language representation learning models (BERT, GPT-2) as well as current state-of-the-art multimodal causality models.

📄 PDF Abstract BibTeX arXiv:2107.10300

Code (0)

등록된 구현이 없습니다.

Tasks

Bias DetectionQuestion AnsweringRepresentation LearningScene UnderstandingVideo CaptioningVideo Question Answering

Similar Papers 제목 키워드 기반

UniReason 1.0: A Unified Reasoning Framework for World Knowledge Aligned Image Generation and Editing

2026-02-02 · Dianyi Wang, Chaofan Ma, Feng Han, Size Wu 외 arxiv

Unified multimodal models often struggle with complex synthesis tasks that demand deep reasoning, and typically treat text-to-image generation and image editing as isolated capabilities rather than interconnected reasoni…

Text-to-Image GenerationImage Editing

iReasoner: Trajectory-Aware Intrinsic Reasoning Supervision for Self-Evolving Large Multimodal Models

2026-01-09 · Meghana Sunil, Manikandarajan Venmathimaran, Muthu Subash Kavitha arxiv

Recent work shows that large multimodal models (LMMs) can self-improve from unlabeled data via self-play and intrinsic feedback. Yet existing self-evolving frameworks mainly reward final outcomes, leaving intermediate re…

Multimodal ReasoningDecision Making

EffiReason-Bench: A Unified Benchmark for Evaluating and Advancing Efficient Reasoning in Large Language Models

2025-11-13 · Junquan Huang, Haotian Wu, Yubo Gao, Yibo Yan 외 arxiv

Large language models (LLMs) with Chain-of-Thought (CoT) prompting achieve strong reasoning but often produce unnecessarily long explanations, increasing cost and sometimes reducing accuracy. Fair comparison of efficienc…

SwiReasoning: Switch-Thinking in Latent and Explicit for Pareto-Superior Reasoning LLMs

2025-10-06 · Dachuan Shi, Abedelkadir Asi, Keying Li, Xiangchi Yuan 외 arxiv

Recent work shows that, beyond discrete reasoning through explicit chain-of-thought steps, which are limited by the boundaries of natural languages, large language models (LLMs) can also reason continuously in latent spa…

Accurate, Interdisciplinary and Transparent Structure-property Understanding with Deep Native Structural Reasoning

2026-07-08 · Chen Tang, Yizhou Wang, Jianyu Wu, Lintao Wang 외 arxiv

Structure-property relationships are foundational to biology, chemistry and materials science, where function, reactivity and physical response emerge from spatial, chemical and periodic organization. Mechanistically exp…

Single-step retrosynthesis