paper-with-me

Papers

How is Visual Attention Influenced by Text Guidance? Database and Model

2024-04-11 · Yinan Sun, Xiongkuo Min, Huiyu Duan, Guangtao Zhai

The analysis and prediction of visual attention have long been crucial tasks in the fields of computer vision and image processing. In practical applications, images are generally accompanied by various text descriptions, however, few studies have explored the influence of text descriptions on visual attention, let alone developed visual saliency prediction models considering text guidance. In this paper, we conduct a comprehensive study on text-guided image saliency (TIS) from both subjective and objective perspectives. Specifically, we construct a TIS database named SJTU-TIS, which includes 1200 text-image pairs and the corresponding collected eye-tracking data. Based on the established SJTU-TIS database, we analyze the influence of various text descriptions on visual attention. Then, to facilitate the development of saliency prediction models considering text influence, we construct a benchmark for the established SJTU-TIS database using state-of-the-art saliency models. Finally, considering the effect of text descriptions on visual attention, while most existing saliency models ignore this impact, we further propose a text-guided saliency (TGSal) prediction model, which extracts and integrates both image features and text features to predict the image saliency under various text-description conditions. Our proposed model significantly outperforms the state-of-the-art saliency models on both the SJTU-TIS database and the pure image saliency databases in terms of various evaluation metrics. The SJTU-TIS database and the code of the proposed TGSal model will be released at: https://github.com/IntMeGroup/TGSal.

📄 PDF Abstract BibTeX arXiv:2404.07537

Code (0)

등록된 구현이 없습니다.

Tasks

PredictionSaliency Prediction

Similar Papers 제목 키워드 기반

Semantic Feature Attention Network for Liver Tumor Segmentation in Large-scale CT database

2019-11-01 · Yao Zhang, Cheng Zhong, Yang Zhang, Zhongchao shi 외

Liver tumor segmentation plays an important role in hepatocellular carcinoma diagnosis and surgical planning. In this paper, we propose a novel Semantic Feature Attention Network (SFAN) for liver tumor segmentation from …

Computed Tomography (CT)SegmentationTumor Segmentation

Focal Guidance: Unlocking Controllability from Semantic-Weak Layers in Video Diffusion Models

2026-01-12 · Yuanyang Yin, Yufan Deng, Shenghai Yuan, Kaipeng Zhang 외 arxiv

The task of Image-to-Video (I2V) generation aims to synthesize a video from a reference image and a text prompt. This requires diffusion models to reconcile high-frequency visual constraints and low-frequency textual gui…

Instruction Following

Attention-space Contrastive Guidance for Efficient Hallucination Mitigation in LVLMs

2026-01-20 · Yujin Jo, Sangyoon Bae, Taesup Kim arxiv

Hallucinations in large vision--language models (LVLMs) often arise when language priors dominate over visual evidence, leading to object misidentification and visually inconsistent descriptions. We address this problem …

Boosting Object Proposals: From Pascal to COCO

2015-12-01 · ICCV 2015 12 · Jordi Pont-Tuset, Luc van Gool

Computer vision in general, and object proposals in particular, are nowadays strongly influenced by the databases on which researchers evaluate the performance of their algorithms. This paper studies the transition from …

Object

Saliency-Guided Attention Network for Image-Sentence Matching

2019-04-20 · ICCV 2019 10 · Zhong Ji, Haoran Wang, Jungong Han, Yanwei Pang

This paper studies the task of matching image and sentence, where learning appropriate representations across the multi-modal data appears to be the main challenge. Unlike previous approaches that predominantly deploy sy…

Sentence