paper-with-me

홈 › Papers

CRIPP-VQA: Counterfactual Reasoning about Implicit Physical Properties via Video Question Answering

2022-11-07 · Maitreya Patel, Tejas Gokhale, Chitta Baral, Yezhou Yang

Videos often capture objects, their visible properties, their motion, and the interactions between different objects. Objects also have physical properties such as mass, which the imaging pipeline is unable to directly capture. However, these properties can be estimated by utilizing cues from relative object motion and the dynamics introduced by collisions. In this paper, we introduce CRIPP-VQA, a new video question answering dataset for reasoning about the implicit physical properties of objects in a scene. CRIPP-VQA contains videos of objects in motion, annotated with questions that involve counterfactual reasoning about the effect of actions, questions about planning in order to reach a goal, and descriptive questions about visible properties of objects. The CRIPP-VQA test set enables evaluation under several out-of-distribution settings -- videos with objects with masses, coefficients of friction, and initial velocities that are not observed in the training distribution. Our experiments reveal a surprising and significant performance gap in terms of answering questions about implicit properties (the focus of this paper) and explicit properties of objects (the focus of prior work).

📄 PDF Abstract BibTeX arXiv:2211.03779

Code (1)

maitreyapatel/cripp-vqa 공식 구현 pytorch

Tasks

Add - POAdd - PQcounterfactualCounterfactual PlanningCounterfactual ReasoningDescriptiveFrictionQuestion AnsweringRemove - PORemove - PQReplace - POReplace - PQVideo Question AnsweringVisual Question Answering (VQA)

Methods 이 논문이 사용한 방법론

Test 설명 없음

Similar Papers 제목 키워드 기반

Disentangled Counterfactual Learning for Physical Audiovisual Commonsense Reasoning

2023-10-30 · NeurIPS 2023 11

In this paper, we propose a Disentangled Counterfactual Learning~(DCL) approach for physical audiovisual commonsense reasoning. The task aims to infer objects' physics commonsense based on both video and audio input, wit…

counterfactual

PASTA: A Dataset for Modeling Participant States in Narratives

2022-07-31 · Sayontan Ghosh, Mahnaz Koupaee, Isabella Chen, Francis Ferraro 외

The events in a narrative are understood as a coherent whole via the underlying states of their participants. Often, these participant states are not explicitly mentioned, instead left to be inferred by the reader. A mod…

BenchmarkingCommon Sense Reasoningcounterfactual

Robust Disentangled Counterfactual Learning for Physical Audiovisual Commonsense Reasoning

2025-02-18 · Mengshi Qi, Changsheng Lv, Huadong Ma

In this paper, we propose a new Robust Disentangled Counterfactual Learning (RDCL) approach for physical audiovisual commonsense reasoning. The task aims to infer objects' physics commonsense based on both video and audi…

counterfactual

Enhancing Physical Plausibility in Video Generation by Reasoning the Implausibility

2025-09-29 · Yutong Hao, Chen Chen, Ajmal Saeed Mian, Chang Xu 외 arxiv

Diffusion models can generate realistic videos, but existing methods rely on implicitly learning physical reasoning from large-scale text-video datasets, which is costly, difficult to scale, and still prone to producing …

Video Generation

Counterfactual Collaborative Reasoning

2023-06-30 · Jianchao Ji, Zelong Li, Shuyuan Xu, Max Xiong 외

Causal reasoning and logical reasoning are two important types of reasoning abilities for human intelligence. However, their relationship has not been extensively explored under machine intelligence context. In this pape…

counterfactualCounterfactual ReasoningData AugmentationLogical Reasoning+1