paper-with-me

Papers

Premise-based Multimodal Reasoning: Conditional Inference on Joint Textual and Visual Clues

2021-05-15 · ACL 2022 5 · Qingxiu Dong, Ziwei Qin, Heming Xia, Tian Feng, Shoujie Tong, Haoran Meng, Lin Xu, Weidong Zhan, Sujian Li, Zhongyu Wei, Tianyu Liu, Zuifang Sui

It is a common practice for recent works in vision language cross-modal reasoning to adopt a binary or multi-choice classification formulation taking as input a set of source image(s) and textual query. In this work, we take a sober look at such an unconditional formulation in the sense that no prior knowledge is specified with respect to the source image(s). Inspired by the designs of both visual commonsense reasoning and natural language inference tasks, we propose a new task termed Premise-based Multi-modal Reasoning(PMR) where a textual premise is the background presumption on each source image. The PMR dataset contains 15,360 manually annotated samples which are created by a multi-phase crowd-sourcing process. With selected high-quality movie screenshots and human-curated premise templates from 6 pre-defined categories, we ask crowd-source workers to write one true hypothesis and three distractors (4 choices) given the premise and image through a cross-check procedure. Besides, we generate adversarial samples to alleviate the annotation artifacts and double the size of PMR. We benchmark various state-of-the-art (pretrained) multi-modal inference models on PMR and conduct comprehensive experimental analyses to showcase the utility of our dataset.

📄 PDF Abstract BibTeX arXiv:2105.07122

Code (0)

등록된 구현이 없습니다.

Tasks

Multimodal ReasoningNatural Language InferenceVisual Commonsense Reasoning

Methods 이 논문이 사용한 방법론

ABC Class of methods in Bayesian Statistics where the posterior distribution is approximated over a rejection scheme on simulations because the likelihood function is…

Similar Papers 제목 키워드 기반

Certain and Uncertain Inference with Indicative Conditionals

2022-07-17 · Paul Égré, Lorenzo Rossi, Jan Sprenger

This paper develops a trivalent semantics for the truth conditions and the probability of the natural language indicative conditional. Our framework rests on trivalent truth conditions first proposed by W. Cooper and yie…

Evaluating BERT for natural language inference: A case study on the CommitmentBank

2019-11-01 · IJCNLP 2019 11 · Nanjiang Jiang, Marie-Catherine de Marneffe

Natural language inference (NLI) datasets (e.g., MultiNLI) were collected by soliciting hypotheses for a given premise from annotators. Such data collection led to annotation artifacts: systems can identify the premise-h…

Natural Language InferenceNegation

VIOLIN: A Large-Scale Dataset for Video-and-Language Inference

2020-03-25 · CVPR 2020 6 · Jingzhou Liu, Wenhu Chen, Yu Cheng, Zhe Gan 외

We introduce a new task, Video-and-Language Inference, for joint multimodal understanding of video and text. Given a video clip with aligned subtitles as premise, paired with a natural language hypothesis based on the vi…

A Multi-Modal Context Reasoning Approach for Conditional Inference on Joint Textual and Visual Clues

2023-05-08 · Yunxin Li, Baotian Hu, Xinyu Chen, Yuxin Ding 외

Conditional inference on joint textual and visual clues is a multi-modal reasoning task that textual clues provide prior permutation or external knowledge, which are complementary with visual content and pivotal to deduc…

Language ModelingLanguage Modelling

PREMISE: Scalable and Strategic Prompt Optimization for Efficient Mathematical Reasoning in Large Models

2025-06-12 · Ye Yu, Yaoning Yu, Haohan Wang

Large reasoning models (LRMs) such as Claude 3.7 Sonnet and OpenAI o1 achieve strong performance on mathematical benchmarks using lengthy chain-of-thought (CoT) reasoning, but the resulting traces are often unnecessarily…

GSM8KMathematical Reasoning