paper-with-me

홈 › Papers

Personalized Image Descriptions from Attention Sequences

2025-12-07 · Ruoyu Xue, Hieu Le, Jingyi Xu, Sounak Mondal, Abe Leite, Gregory Zelinsky, Minh Hoai, Dimitris Samaras arxiv

People can view the same image differently: they focus on different regions, objects, and details in varying orders and describe them in distinct linguistic styles. This leads to substantial variability in image descriptions. However, existing models for personalized image description focus on linguistic style alone, with no prior work leveraging individual viewing patterns. We address this gap by explicitly modeling personalized viewing behavior as a core factor in description generation. Our method, DEPER (DEscription-PERception persona encoder), learns a subject embedding that captures both linguistic style and viewing behavior, guided by an auxiliary attention-prediction task. A lightweight adapter aligns these embeddings with a frozen vision-language model, enabling few-shot personalization without retraining. Across four datasets spanning diverse viewing tasks and both short and detailed descriptions, DEPER achieves a 24% average improvement, showing that modeling personalized attention produces more human-aligned and high-quality descriptions. We posit that understanding how people see helps predict what they say; modeling human diversity in perception can improve both performance and human alignment in multimodal systems.

📄 PDF Abstract BibTeX arXiv:2512.06662

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Every Image Listens, Every Image Dances: Music-Driven Image Animation

2025-01-30 · Zhikang Dong, Weituo Hao, Ju-Chiang Wang, Peng Zhang 외

Image animation has become a promising area in multimodal research, with a focus on generating videos from reference images. While prior work has largely emphasized generic video generation guided by text, music-driven d…

Image AnimationVideo Generation

Subject-Diffusion:Open Domain Personalized Text-to-Image Generation without Test-time Fine-tuning

2023-07-21 · Jian Ma, Junhao Liang, Chen Chen, Haonan Lu

Recent progress in personalized image generation using diffusion models has been significant. However, development in the area of open-domain and non-fine-tuning personalized image generation is proceeding rather slowly.…

Diffusion PersonalizationDiffusion Personalization Tuning FreeImage GenerationPersonalized Image Generation+2

Conceptrol: Concept Control of Zero-shot Personalized Image Generation

2025-03-09 · Qiyuan He, Angela Yao

Personalized image generation with text-to-image diffusion models generates unseen images based on reference image content. Zero-shot adapter methods such as IP-Adapter and OminiControl are especially interesting because…

Image GenerationPersonalized Image Generation

Personalization of Saliency Estimation

2017-11-21 · Bingqing Yu, James J. Clark

Most existing saliency models use low-level features or task descriptions when generating attention predictions. However, the link between observer characteristics and gaze patterns is rarely investigated. We present a n…

Saliency Prediction

Egocentric Video Description based on Temporally-Linked Sequences

2017-04-07 · Marc Bolaños, Álvaro Peris, Francisco Casacuberta, Sergi Soler 외

Egocentric vision consists in acquiring images along the day from a first person point-of-view using wearable cameras. The automatic analysis of this information allows to discover daily patterns for improving the qualit…

DecoderVideo Description