paper-with-me

홈 › Papers

Small-Large Collaboration: Training-efficient Concept Personalization for Large VLM using a Meta Personalized Small VLM

2025-08-10 · Sihan Yang, Huitong Ji, Shaolin Lu, Jiayi Chen, Binxiao Xu, Ming Lu, Yuanxing Zhang, Wenhui Dong, Wentao Zhang arxiv

Personalizing Vision-Language Models (VLMs) to transform them into daily assistants has emerged as a trending research direction. However, leading companies like OpenAI continue to increase model size and develop complex designs such as the chain of thought (CoT). While large VLMs are proficient in complex multi-modal understanding, their high training costs and limited access via paid APIs restrict direct personalization. Conversely, small VLMs are easily personalized and freely available, but they lack sufficient reasoning capabilities. Inspired by this, we propose a novel collaborative framework named Small-Large Collaboration (SLC) for large VLM personalization, where the small VLM is responsible for generating personalized information, while the large model integrates this personalized information to deliver accurate responses. To effectively incorporate personalized information, we develop a test-time reflection strategy, preventing the potential hallucination of the small VLM. Since SLC only needs to train a meta personalized small VLM for the large VLMs, the overall process is training-efficient. To the best of our knowledge, this is the first training-efficient framework that supports both open-source and closed-source large VLMs, enabling broader real-world personalized applications. We conduct thorough experiments across various benchmarks and large VLMs to demonstrate the effectiveness of the proposed SLC framework. The code will be released at https://github.com/Hhankyangg/SLC.

📄 PDF Abstract BibTeX arXiv:2508.07260

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Enabling Personalized Long-term Interactions in LLM-based Agents through Persistent Memory and User Profiles

2025-10-09 · Rebecca Westhäußer, Wolfgang Minker, Sebatian Zepf arxiv

Large language models (LLMs) increasingly serve as the central control unit of AI agents, yet current approaches remain limited in their ability to deliver personalized interactions. While Retrieval Augmented Generation …

Key-Locked Rank One Editing for Text-to-Image Personalization

2023-05-02 · Yoad Tewel, Rinon Gal, Gal Chechik, Yuval Atzmon

Text-to-image models (T2I) offer a new level of flexibility by allowing users to guide the creative process through natural language. However, personalizing these models to align with user-provided visual concepts remain…

Mod-Adapter: Tuning-Free and Versatile Multi-concept Personalization via Modulation Adapter

2025-05-24 · Weizhi Zhong, Huan Yang, Zheng Liu, Huiguo He 외

Personalized text-to-image generation aims to synthesize images of user-provided concepts in diverse contexts. Despite recent progress in multi-concept personalization, most are limited to object concepts and struggle to…

Image GenerationMixture-of-ExpertsText to Image GenerationText-to-Image Generation

Is This Loss Informative? Faster Text-to-Image Customization by Tracking Objective Dynamics

2023-02-09 · NeurIPS 2023 11 · Anton Voronov, Mikhail Khoroshikh, Artem Babenko, Max Ryabinin

Text-to-image generation models represent the next step of evolution in image synthesis, offering a natural way to achieve flexible yet fine-grained control over the result. One emerging area of research is the fast adap…

GPUImage GenerationText to Image GenerationText-to-Image Generation

Encoder-based Domain Tuning for Fast Personalization of Text-to-Image Models

2023-02-23 · Rinon Gal, Moab Arar, Yuval Atzmon, Amit H. Bermano 외

Text-to-image personalization aims to teach a pre-trained diffusion model to reason about novel, user provided concepts, embedding them into new scenes guided by natural language prompts. However, current personalization…

Novel Concepts