paper-with-me

홈 › Papers

Ego: Embedding-Guided Personalization of Vision-Language Models

2026-03-10 · Soroush Seifi, Simon Gardier, Vaggelis Dorovatas, Daniel Olmeda Reino, Rahaf Aljundi arxiv

AI assistants that support humans in daily life are becoming increasingly feasible, driven by the rapid advancements in multimodal language models. A key challenge lies in overcoming the generic nature of these models to deliver personalized experiences. Existing approaches to personalizing large vision language models often rely on additional training stages, which limit generality and scalability, or on engineered pipelines with external pre-trained modules, which hinder deployment efficiency. In this work, we propose an efficient personalization method that leverages the model's inherent ability to capture personalized concepts. Specifically, we extract visual tokens that predominantly represent the target concept by utilizing the model's internal attention mechanisms. These tokens serve as a memory of that specific concept, enabling the model to recall and describe it when it appears in test images. We conduct a comprehensive and unified evaluation of our approach and SOTA methods across various personalization settings including single-concept, multi-concept, and video personalization, demonstrating strong performance gains with minimal personalization overhead.

📄 PDF Abstract BibTeX arXiv:2603.09771

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Align 3D Representation and Text Embedding for 3D Content Personalization

2025-08-23 · Qi Song, Ziyuan Luo, Ka Chun Cheung, Simon See 외 arxiv

Recent advances in NeRF and 3DGS have significantly enhanced the efficiency and quality of 3D content synthesis. However, efficient personalization of generated 3D content remains a critical challenge. Current 3D persona…

Knowledge Distillation

Meta-Personalizing Vision-Language Models to Find Named Instances in Video

2023-06-16 · CVPR 2023 1 · Chun-Hsiao Yeh, Bryan Russell, Josef Sivic, Fabian Caba Heilbron 외

Large-scale vision-language models (VLM) have shown impressive results for language-guided search applications. While these models allow category-level queries, they currently struggle with personalized searches for mome…

RetrievalWord Embeddings

GPFedRec: Graph-guided Personalization for Federated Recommendation

2023-05-13 · Chunxu Zhang, Guodong Long, Tianyi Zhou, Zijjian Zhang 외

The federated recommendation system is an emerging AI service architecture that provides recommendation services in a privacy-preserving manner. Using user-relation graphs to enhance federated recommendations is a promis…

Federated LearningPrivacy PreservingRelation

Encoder-based Domain Tuning for Fast Personalization of Text-to-Image Models

2023-02-23 · Rinon Gal, Moab Arar, Yuval Atzmon, Amit H. Bermano 외

Text-to-image personalization aims to teach a pre-trained diffusion model to reason about novel, user provided concepts, embedding them into new scenes guided by natural language prompts. However, current personalization…

Novel Concepts

Surgeon Style Fingerprinting and Privacy Risk Quantification via Discrete Diffusion Models in a Vision-Language-Action Framework

2025-06-09 · Huixin Zhan, Jason H. Moore

Surgeons exhibit distinct operating styles due to differences in training, experience, and motor behavior - yet current AI systems often ignore this personalization signal. We propose a novel approach to model fine-grain…

DenoisingVision-Language-Action