paper-with-me

Papers

PRIMA: Multi-Image Vision-Language Models for Reasoning Segmentation

2024-12-19 · Muntasir Wahed, Kiet A. Nguyen, Adheesh Sunil Juvekar, Xinzhuo Li, Xiaona Zhou, Vedant Shah, Tianjiao Yu, Pinar Yanardag, Ismini Lourentzou

Despite significant advancements in Large Vision-Language Models (LVLMs), existing pixel-grounding models operate on single-image settings, limiting their ability to perform detailed, fine-grained comparisons across multiple images. Conversely, current multi-image understanding models lack pixel-level grounding. Our work addresses this gap by introducing the task of multi-image pixel-grounded reasoning segmentation, and PRIMA, a novel LVLM that integrates pixel-level grounding with robust multi-image reasoning capabilities to produce contextually rich, pixel-grounded explanations. Central to PRIMA is an efficient vision module that queries fine-grained visual representations across multiple images, reducing TFLOPs by $25.3\%$. To support training and evaluation, we curate $M^4Seg$, a new reasoning segmentation benchmark consisting of $\sim$224K question-answer pairs that require fine-grained visual understanding across multiple images. Experimental results demonstrate PRIMA outperforms state-of-the-art baselines.

📄 PDF Abstract BibTeX arXiv:2412.15209

Code (0)

등록된 구현이 없습니다.

Tasks

Reasoning Segmentation

Similar Papers 제목 키워드 기반

CheXthought: A global multimodal dataset of clinical chain-of-thought reasoning and visual attention for chest X-ray interpretation

2026-04-29 · Sonali Sharma, Jin Long, George Shih, Sarah Eid 외 arxiv

Chest X-ray interpretation is one of the most frequently performed diagnostic tasks in medicine and a primary target for AI development, yet current vision-language models are primarily trained on datasets of paired imag…

Multi-modal Latent Space Learning for Chain-of-Thought Reasoning in Language Models

2023-12-14 · Liqi He, Zuchao Li, Xiantao Cai, Ping Wang

Chain-of-thought (CoT) reasoning has exhibited impressive performance in language models for solving complex tasks and answering questions. However, many real-world questions require multi-modal information, such as text…

Machine Translation

Simple Vision-Language Math Reasoning via Rendered Text

2025-11-12 · Matvey Skripkin, Elizaveta Goncharova, Andrey Kuznetsov arxiv

We present a lightweight yet effective pipeline for training vision-language models to solve math problems by rendering LaTeX encoded equations into images and pairing them with structured chain-of-thought prompts. This …

Chitrarth: Bridging Vision and Language for a Billion People

2025-02-21 · Shaharukh Khan, Ayush Tarun, Abhinav Ravi, Ali Faraz 외

Recent multimodal foundation models are primarily trained on English or high resource European language data, which hinders their applicability to other medium and low-resource languages. To address this limitation, we i…

DiversityLanguage ModelingLanguage ModellingLarge Language Model+1

Stop Pre-Training: Adapt Visual-Language Models to Unseen Languages

2023-06-29 · Yasmine Karoui, Rémi Lebret, Negar Foroutan, Karl Aberer

Vision-Language Pre-training (VLP) has advanced the performance of many vision-language tasks, such as image-text retrieval, visual entailment, and visual reasoning. The pre-training mostly utilizes lexical databases and…

Image-text RetrievalMachine TranslationRetrievalText Retrieval+2