paper-with-me

홈 › Papers

Multi-Task Learning for Visually Grounded Reasoning in Gastrointestinal VQA

2025-11-06 · Itbaan Safwan, Muhammad Annas Shaikh, Muhammad Haaris, Ramail Khan, Muhammad Atif Tahir arxiv

We present a multi-task framework for the MediaEval Medico 2025 challenge, leveraging a LoRA-tuned Florence-2 model for simultaneous visual question answering (VQA), explanation generation, and visual grounding. The proposed system integrates three curated datasets: (1) Kvasir-VQA-x1 for question-answer learning, (2) a synthetically enriched explanation dataset offering structured medical reasoning, and (3) text-to-region pairs linking visual features with segmentation masks. This multi-task setup enables the model to jointly learn visual grounding, reasoning, and interpretation, producing responses that are both accurate and interpretable. Extensive evaluation demonstrates that our approach substantially improves over single-task baselines in both answer accuracy and visual localization, highlighting the effectiveness of grounded multi-task learning for medical VQA applications.

📄 PDF Abstract BibTeX arXiv:2511.04384

Code (0)

등록된 구현이 없습니다.

Tasks

Visual Question AnsweringExplanation GenerationMulti-Task LearningVisual Localization

Similar Papers 제목 키워드 기반

VoCoT: Unleashing Visually Grounded Multi-Step Reasoning in Large Multi-Modal Models

2024-05-27 · Zejun Li, Ruipu Luo, Jiwen Zhang, Minghui Qiu 외

While large multi-modal models (LMMs) have exhibited impressive capabilities across diverse tasks, their effectiveness in handling complex tasks has been limited by the prevailing single-step reasoning paradigm. To this …

Object

ReasonCLIP-58M: Visually Grounded Commonsense Reasoning Supervision for CLIP

2026-06-25 · Sicheng Zhang, Muzammal Naseer, Binzhu Xie, Naufal Suryanto 외 arxiv

CLIP and its variants are widely adopted visual backbones in multimodal systems, but their pretraining remains dominated by descriptive image-text alignment. As downstream applications increasingly demand visually ground…

Continual Pretraining

Affective Visual Dialog: A Large-Scale Benchmark for Emotional Reasoning Based on Visually Grounded Conversations

2023-08-30 · Kilichbek Haydarov, Xiaoqian Shen, Avinash Madasu, Mahmoud Salem 외

We introduce Affective Visual Dialog, an emotion explanation and reasoning task as a testbed for research on understanding the formation of emotions in visually grounded conversations. The task involves three skills: (1)…

Explanation GenerationQuestion AnsweringVisual Dialog

iVGR: Internalizing Visually Grounded Reasoning for MLLMs with Reinforcement Learning

2026-05-29 · Chang-Bin Zhang, Yujie Zhong, Qiang Zhang, Kai Han arxiv

While visually grounded Chain-of-Thought (CoT) has emerged as a promising paradigm to enhance fine-grained perception in multimodal large language models (MLLMs), its efficacy during the inference phase remains underexpl…

Reinforcement LearningVisual LocalizationVisual Grounding

Point-RFT: Improving Multimodal Reasoning with Visually Grounded Reinforcement Finetuning

2025-05-26 · Minheng Ni, Zhengyuan Yang, Linjie Li, Chung-Ching Lin 외

Recent advances in large language models have significantly improved textual reasoning through the effective use of Chain-of-Thought (CoT) and reinforcement learning. However, extending these successes to vision-language…

document understandingMultimodal ReasoningVisual Reasoning