paper-with-me

홈 › Papers

Through the Magnifying Glass: Adaptive Perception Magnification for Hallucination-Free VLM Decoding

2025-03-13 · Shunqi Mao, Chaoyi Zhang, Weidong Cai

Existing vision-language models (VLMs) often suffer from visual hallucination, where the generated responses contain inaccuracies that are not grounded in the visual input. Efforts to address this issue without model finetuning primarily mitigate hallucination by reducing biases contrastively or amplifying the weights of visual embedding during decoding. However, these approaches improve visual perception at the cost of impairing the language reasoning capability. In this work, we propose the Perception Magnifier (PM), a novel visual decoding method that iteratively isolates relevant visual tokens based on attention and magnifies the corresponding regions, spurring the model to concentrate on fine-grained visual details during decoding. Specifically, by magnifying critical regions while preserving the structural and contextual information at each decoding step, PM allows the VLM to enhance its scrutiny of the visual input, hence producing more accurate and faithful responses. Extensive experimental results demonstrate that PM not only achieves superior hallucination mitigation but also enhances language generation while preserving strong reasoning capabilities.Code is available at https://github.com/ShunqiM/PM .

📄 PDF Abstract BibTeX arXiv:2503.10183

Code (0)

등록된 구현이 없습니다.

Tasks

HallucinationText Generation

Methods 이 논문이 사용한 방법론

Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
Attention 설명 없음

Similar Papers 제목 키워드 기반

Event-Based Motion Magnification

2024-02-19 · Yutian Chen, Shi Guo, Fangzheng Yu, Feng Zhang 외

Detecting and magnifying imperceptible high-frequency motions in real-world scenarios has substantial implications for industrial and medical applications. These motions are characterized by small amplitudes and high fre…

BenchmarkingMotion DetectionMotion Magnification

Self-Supervised Motion Magnification by Backpropagating Through Optical Flow

2023-11-28 · NeurIPS 2023 11

This paper presents a simple, self-supervised method for magnifying subtle motions in video: given an input video and a magnification factor, we manipulate the video such that its new optical flow is scaled by the desire…

Motion MagnificationOptical Flow EstimationTest-time Adaptation

DeepMag: Source Specific Motion Magnification Using Gradient Ascent

2018-08-09 · Weixuan Chen, Daniel McDuff

Many important physical phenomena involve subtle signals that are difficult to observe with the unaided eye, yet visualizing them can be very informative. Current motion magnification techniques can reveal these small te…

Motion Magnification

Unsupervised Behaviour Analysis and Magnification (uBAM) using Deep Learning

2020-12-16 · Biagio Brattoli, Uta Buechler, Michael Dorkenwald, Philipp Reiser 외

Motor behaviour analysis is essential to biomedical research and clinical diagnostics as it provides a non-invasive strategy for identifying motor impairment and its change caused by interventions. State-of-the-art instr…

Deep LearningDiagnostic

Magnifying Subtle Facial Motions for Effective 4D Expression Recognition

2021-05-05 · Qingkai Zhen, Di Huang, Yunhong Wang, Hassen Drira 외

In this paper, an effective pipeline to automatic 4D Facial Expression Recognition (4D FER) is proposed. It combines two growing but disparate ideas in Computer Vision -- computing the spatial facial deformations using t…

Emotion ClassificationFacial Expression RecognitionFacial Expression Recognition (FER)