paper-with-me

홈 › Papers

Patch Matters: Training-free Fine-grained Image Caption Enhancement via Local Perception

2025-01-01 · CVPR 2025 1 · Ruotian Peng, Haiying He, Yake Wei, Yandong Wen, Di Hu

High-quality image captions play a crucial role in improving the performance of cross-modal applications such as text-to-image generation, text-to-video generation, and text-image retrieval. To generate long-form, high-quality captions, many recent studies have employed multimodal large language models (MLLMs). However, current MLLMs often produce captions that lack fine-grained details or suffer from hallucinations, a challenge that persists in both open-source and closed-source models. Inspired by Feature-Integration theory, which suggests that attention must focus on specific regions to integrate visual information effectively, we propose a divide-then-aggregate strategy. Our method first divides the image into semantic and spatial patches to extract fine-grained details, enhancing the model's local perception of the image. These local details are then hierarchically aggregated to generate a comprehensive global description. To address hallucinations and inconsistencies in the generated captions, we apply a semantic-level filtering process during hierarchical aggregation. This training-free pipeline can be applied to both open-source models (LLaVA-1.5, LLaVA-1.6, Mini-Gemini) and closed-source models (Claude-3.5-Sonnet, GPT-4o, GLM-4V-Plus). Extensive experiments demonstrate that our method generates more detailed, reliable captions, advancing multimodal description generation without requiring model retraining.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Image CaptioningImage GenerationImage RetrievalText to Image GenerationText-to-Image GenerationText-to-Video GenerationVideo Generation

Methods 이 논문이 사용한 방법론

Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
Attention 설명 없음
Focus 설명 없음

Similar Papers 제목 키워드 기반

Look Where It Matters: Training-Free Ultra-HR Remote Sensing VQA via Adaptive Zoom Search

2025-11-25 · Yunqi Zhou, Chengjie Jiang, Chun Yuan, Jing Li arxiv

With advances in satellite constellations, sensor technologies, and imaging pipelines, ultra-high-resolution (Ultra-HR) remote sensing imagery is becoming increasingly widespread. However, current remote sensing foundati…

Visual Question Answering

Dynamic Granularity Matters: Rethinking Vision Transformers Beyond Fixed Patch Splitting

2025-11-24 · Qiyang Yu, Yu Fang, Tianrui Li, Xuemei Cao 외 arxiv

Vision Transformers (ViTs) have demonstrated strong capabilities in capturing global dependencies but often struggle to efficiently represent fine-grained local details. Existing multi-scale approaches alleviate this iss…

Computational Efficiency

Mask Guided Attention For Fine-Grained Patchy Image Classification

2021-02-04 · Jun Wang, Xiaohan Yu, Yongsheng Gao

In this work, we present a novel mask guided attention (MGA) method for fine-grained patchy image classification. The key challenge of fine-grained patchy image classification lies in two folds, ultra-fine-grained inter-…

ClassificationGeneral Classificationimage-classificationImage Classification+2

When Polysemy Matters: Modeling Semantic Categorization with Word Embeddings

2022-07-01 · *SEM (NAACL) 2022 7 · Elizabeth Soper, Jean-Pierre Koenig

Recent work using word embeddings to model semantic categorization have indicated that static models outperform the more recent contextual class of models (Majewska et al, 2021). In this paper, we consider polysemy as a …

Word Embeddings

Shape-aware Sampling Matters in the Modeling of Multi-Class Tubular Structures

2025-06-14 · Minghui Zhang, Yaoyu Liu, Xin You, Hanxiao Zhang 외

Accurate multi-class tubular modeling is critical for precise lesion localization and optimal treatment planning. Deep learning methods enable automated shape modeling by prioritizing volumetric overlap accuracy. However…