paper-with-me

홈 › Papers

Pensieve: Retrospect-then-Compare Mitigates Visual Hallucination

2024-03-21 · Dingchen Yang, Bowen Cao, Guang Chen, Changjun Jiang

Multi-modal Large Language Models (MLLMs) demonstrate remarkable success across various vision-language tasks. However, they suffer from visual hallucination, where the generated responses diverge from the provided image. Are MLLMs oblivious to the accurate visual cues when they hallucinate? Our investigation reveals that the visual branch may equally advocate both accurate and erroneous content. To address this issue, we propose Pensieve, a training-free method that leverages the analogous visual hallucinations, which are induced by images sharing common semantic and appearance characteristics, to mitigate hallucination. Specifically, Pensieve enables MLLMs to retrospect relevant images as references and compare their visual content with the test image via confidence score subtraction. Moreover, our paradigm balances the effects of addressing errors from both the visual and textual branches by adaptively scaling the subtracted scores. Experiments on Whoops, LLaVA Bench, POPE, and MME demonstrate the efficacy of Pensieve in mitigating visual hallucination, surpassing other advanced decoding strategies. Pensieve also aids MLLMs in identifying visual details and enhance the specificity of generated image descriptions.

📄 PDF Abstract BibTeX arXiv:2403.14401

Code (1)

dingchenyang99/pensieve 공식 구현 pytorch

Tasks

HallucinationMMESpecificity

Similar Papers 제목 키워드 기반

Pensieve 5G: Implementation of RL-based ABR Algorithm for UHD 4K/8K Content Delivery on Commercial 5G SA/NR-DC Network

2022-12-29 · Kasidis Arunruangsirilert, Bo Wei, Hang Song, Jiro Katto

While the rollout of the fifth-generation mobile network (5G) is underway across the globe with the intention to deliver 4K/8K UHD videos, Augmented Reality (AR), and Virtual Reality (VR) content to the mass amounts of u…

4k8k

Stateful Large Language Model Serving with Pensieve

2023-12-09 · Lingfan Yu, JinKun Lin, Jinyang Li

Large Language Models (LLMs) are wildly popular today and it is important to serve them efficiently. Existing LLM serving systems are stateless across requests. Consequently, when LLMs are used in the common setting of m…

CPUGPULanguage ModelingLanguage Modelling+2

Pensieve Grader: An AI-Powered, Ready-to-Use Platform for Effortless Handwritten STEM Grading

2025-07-02 · Yoonseok Yang, Minjune Kim, Marlon Rondinelli, Keren Shao arxiv

Grading handwritten, open-ended responses remains a major bottleneck in large university STEM courses. We introduce Pensieve (https://www.pensieve.co), an AI-assisted grading platform that leverages large language models…

Memory-QA: Answering Recall Questions Based on Multimodal Memories

2025-09-22 · Hongda Jiang, Xinyuan Zhang, Siddhant Garg, Rishab Arora 외 arxiv

We introduce Memory-QA, a novel real-world task that involves answering recall questions about visual content from previously stored multimodal memories. This task poses unique challenges, including the creation of task-…

NANCY: Neural Adaptive Network Coding methodologY for video distribution over wireless networks

2020-08-21 · Paresh Saxena, Mandan Naresh, Manik Gupta, Anirudh Achanta 외

This paper presents NANCY, a system that generates adaptive bit rates (ABR) for video and adaptive network coding rates (ANCR) using reinforcement learning (RL) for video distribution over wireless networks. NANCY trains…

reinforcement-learningReinforcement Learning (RL)