paper-with-me

홈 › Papers

POPEN: Preference-Based Optimization and Ensemble for LVLM-Based Reasoning Segmentation

2025-01-01 · CVPR 2025 1 · Lanyun Zhu, Tianrun Chen, Qianxiong Xu, Xuanyi Liu, Deyi Ji, Haiyang Wu, De Wen Soh, Jun Liu

Existing LVLM-based reasoning segmentation methods often suffer from imprecise segmentation results and hallucinations in their text responses. This paper introduces POPEN, a novel framework designed to address these issues and achieve improved results. POPEN includes a preference-based optimization method to finetune the LVLM, aligning it more closely with human preferences and thereby generating better text responses and segmentation results. Additionally, POPEN introduces a preference-based ensemble method for inference, which integrates multiple outputs from the LVLM using a preference-score-based attention mechanism for refinement. To better adapt to the segmentation task, we incorporate several task-specific designs in our POPEN framework, including a new approach for collecting segmentation preference data with a curriculum learning mechanism, and a novel preference optimization loss to refine the segmentation capability of the LVLM. Experiments demonstrate that our method achieves state-of-the-art performance in reasoning segmentation, exhibiting minimal hallucination in text responses and the highest segmentation accuracy compared to previous advanced methods like LISA and PixelLM.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

HallucinationReasoning SegmentationSegmentation

Methods 이 논문이 사용한 방법론

Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
Attention 설명 없음

Similar Papers 제목 키워드 기반

MedAlign: A Synergistic Framework of Multimodal Preference Optimization and Federated Meta-Cognitive Reasoning

2025-10-24 · Siyong Chen, Jinbo Wen, Jiawen Kang, Tenghui Huang 외 arxiv

Recently, large models have shown significant potential for smart healthcare. However, the deployment of Large Vision-Language Models (LVLMs) for clinical services is currently hindered by three critical challenges: a te…

Visual Question Answering

Improving Generalization in Visual Reasoning via Self-Ensemble

2024-10-28 · Tien-Huy Nguyen, Quang-Khai Tran, Anh-Tuan Quang-Hoang

The cognitive faculty of visual reasoning necessitates the integration of multimodal perceptual processing and commonsense and external knowledge of the world. In recent years, a plethora of large vision-language models …

Visual Question Answering (VQA)Visual Reasoning

MMedPO: Aligning Medical Vision-Language Models with Clinical-Aware Multimodal Preference Optimization

2024-12-09 · Kangyu Zhu, Peng Xia, Yun Li, Hongtu Zhu 외

The advancement of Large Vision-Language Models (LVLMs) has propelled their application in the medical field. However, Medical LVLMs (Med-LVLMs) encounter factuality challenges due to modality misalignment, where the mod…

Visual Question Answering (VQA)

Self-alignment of Large Video Language Models with Refined Regularized Preference Optimization

2025-04-16 · Pritam Sarkar, Ali Etemad

Despite recent advances in Large Video Language Models (LVLMs), they still struggle with fine-grained temporal understanding, hallucinate, and often make simple mistakes on even simple video question-answering tasks, all…

HallucinationQuestion AnsweringVideo Question AnsweringVideo Understanding

VaPR -- Vision-language Preference alignment for Reasoning

2025-10-02 · Rohan Wadhawan, Fabrice Y Harel-Canada, Zi-Yi Dou, Suhaila Shakiah 외 arxiv

Preference finetuning methods like Direct Preference Optimization (DPO) with AI-generated feedback have shown promise in aligning Large Vision-Language Models (LVLMs) with human preferences. However, existing techniques …

Response Generation