paper-with-me

Papers

MMRefine: Unveiling the Obstacles to Robust Refinement in Multimodal Large Language Models

2025-06-05 · Gio Paik, Geewook Kim, Jinbae Im

This paper introduces MMRefine, a MultiModal Refinement benchmark designed to evaluate the error refinement capabilities of Multimodal Large Language Models (MLLMs). As the emphasis shifts toward enhancing reasoning during inference, MMRefine provides a framework that evaluates MLLMs' abilities to detect and correct errors across six distinct scenarios beyond just comparing final accuracy before and after refinement. Furthermore, the benchmark analyzes the refinement performance by categorizing errors into six error types. Experiments with various open and closed MLLMs reveal bottlenecks and factors impeding refinement performance, highlighting areas for improvement in effective reasoning enhancement. Our code and dataset are publicly available at https://github.com/naver-ai/MMRefine.

📄 PDF Abstract BibTeX arXiv:2506.04688

Code (1)

naver-ai/mmrefine 공식 구현

Similar Papers 제목 키워드 기반

Unveiling the Potential of Multimodal Retrieval Augmented Generation with Planning

2025-01-26 · Xiaohan Yu, Zhihan Yang, Chong Chen

Multimodal Retrieval Augmented Generation (MRAG) systems, while promising for enhancing Multimodal Large Language Models (MLLMs), often rely on rigid, single-step retrieval methods. This limitation hinders their ability …

RetrievalRetrieval-augmented Generation

Zero-Shot, But at What Cost? Unveiling the Hidden Overhead of MILS's LLM-CLIP Framework for Image Captioning

2025-04-21 · Yassir Benhammou, Alessandro Tiberio, Gabriel Trautmann, Suman Kalyan

MILS (Multimodal Iterative LLM Solver) is a recently published framework that claims "LLMs can see and hear without any training" by leveraging an iterative, LLM-CLIP based approach for zero-shot image captioning. While …

Image Captioning

Guided Query Refinement: Multimodal Hybrid Retrieval with Test-Time Optimization

2025-10-06 · Omri Uzan, Asaf Yehudai, Roi pony, Eyal Shnarch 외 arxiv

Multimodal encoders have pushed the boundaries of visual document retrieval, matching textual query tokens directly to image patches and achieving state-of-the-art performance on public benchmarks. Recent models relying …

Unveiling the Tapestry of Consistency in Large Vision-Language Models

2024-05-23 · Yuan Zhang, Fei Xiao, Tao Huang, Chun-Kai Fan 외

Large vision-language models (LVLMs) have recently achieved rapid progress, exhibiting great perception and reasoning abilities concerning visual information. However, when faced with prompts in different sizes of soluti…

Diagnostic

Obstacle Identification and Ellipsoidal Decomposition for Fast Motion Planning in Unknown Dynamic Environments

2022-09-28 · Mehmetcan Kaymaz, Nazim Kemal Ure

Collision avoidance in the presence of dynamic obstacles in unknown environments is one of the most critical challenges for unmanned systems. In this paper, we present a method that identifies obstacles in terms of ellip…

Collision AvoidanceMotion Planning