paper-with-me

Papers

Exploring Failure Cases in Multimodal Reasoning About Physical Dynamics

2024-02-24 · Sadaf Ghaffari, Nikhil Krishnaswamy

In this paper, we present an exploration of LLMs' abilities to problem solve with physical reasoning in situated environments. We construct a simple simulated environment and demonstrate examples of where, in a zero-shot setting, both text and multimodal LLMs display atomic world knowledge about various objects but fail to compose this knowledge in correct solutions for an object manipulation and placement task. We also use BLIP, a vision-language model trained with more sophisticated cross-modal attention, to identify cases relevant to object physical properties that that model fails to ground. Finally, we present a procedure for discovering the relevant properties of objects in the environment and propose a method to distill this knowledge back into the LLM.

📄 PDF Abstract BibTeX arXiv:2402.15654

Code (0)

등록된 구현이 없습니다.

Tasks

Language ModelingLanguage ModellingMultimodal ReasoningObjectWorld Knowledge

Methods 이 논문이 사용한 방법론

BLIP Vision-Language Pre-training (VLP) has advanced the performance for many vision-language tasks. However, most existing pre-trained models only excel in either understanding-based…

Similar Papers 제목 키워드 기반

Exploring Diagnostic Prompting Approach for Multimodal LLM-based Visual Complexity Assessment: A Case Study of Amazon Search Result Pages

2025-11-26 · Divendar Murtadak, Yoon Kim, Trilokya Akula arxiv

This study investigates whether diagnostic prompting can improve Multimodal Large Language Model (MLLM) reliability for visual complexity assessment of Amazon Search Results Pages (SRP). We compare diagnostic prompting w…

Evaluating Multi-Turn Multimodal Diagnostic Reasoning on Challenging Real-World Clinical Cases

2026-07-28 · Rui Yang, Weihao Xuan, Yi Lin, Zhuhan Bao 외 arxiv

Clinical diagnostic evaluation should not only assess whether models can provide correct diagnoses, but also reflect the realities of clinical practice, including progressive disclosure of multimodal information, dynamic…

MINERVA: Evaluating Complex Video Reasoning

2025-05-01 · Arsha Nagrani, Sachit Menon, Ahmet Iscen, Shyamal Buch 외

Multimodal LLMs are turning their focus to video benchmarks, however most video benchmarks only provide outcome supervision, with no intermediate or interpretable reasoning steps. This makes it challenging to assess if m…

BenchmarkingTemporal Localization

Evaluating Self-Correcting Vision Agents Through Quantitative and Qualitative Metrics

2026-01-14 · Aradhya Dixit arxiv

Recent progress in multimodal foundation models has enabled Vision-Language Agents (VLAs) to decompose complex visual tasks into executable tool-based plans. While recent benchmarks have begun to evaluate iterative self-…

Exploring the Limitations of Large Language Models in Compositional Relation Reasoning

2024-03-05 · Jinman Zhao, Xueyan Zhang

We present a comprehensive evaluation of large language models(LLMs)' ability to reason about composition relations through a benchmark encompassing 1,500 test cases in English, designed to cover six distinct types of co…

Relation