paper-with-me

Papers

Personalized Vision via Visual In-Context Learning

2025-09-29 · Yuxin Jiang, Yuchao Gu, Yiren Song, Ivor Tsang, Mike Zheng Shou arxiv

Modern vision models, trained on large-scale annotated datasets, excel at predefined tasks but struggle with personalized vision -- tasks defined at test time by users with customized objects or novel objectives. Existing personalization approaches rely on costly fine-tuning or synthetic data pipelines, which are inflexible and restricted to fixed task formats. Visual in-context learning (ICL) offers a promising alternative, yet prior methods confine to narrow, in-domain tasks and fail to generalize to open-ended personalization. We introduce Personalized In-Context Operator (PICO), a simple four-panel framework that repurposes diffusion transformers as visual in-context learners. Given a single annotated exemplar, PICO infers the underlying transformation and applies it to new inputs without retraining. To enable this, we construct VisRel, a compact yet diverse tuning dataset, showing that task diversity, rather than scale, drives robust generalization. We further propose an attention-guided seed scorer that improves reliability via efficient inference scaling. Extensive experiments demonstrate that PICO (i) surpasses fine-tuning and synthetic-data baselines, (ii) flexibly adapts to novel user-defined tasks, and (iii) generalizes across both recognition and generation.

📄 PDF Abstract BibTeX arXiv:2509.25172

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Contextualized Visual Personalization in Vision-Language Models

2026-02-03 · Yeongtak Oh, Sangwon Yu, Junsung Park, Han Cheol Moon 외 arxiv

Despite recent progress in vision-language models (VLMs), existing approaches often fail to generate personalized responses based on the user's specific experiences, as they lack the ability to associate visual inputs wi…

Image Captioning

Teaching VLMs to Localize Specific Objects from In-context Examples

2024-11-20 · Sivan Doveh, Nimrod Shabtay, Wei Lin, Eli Schwartz 외

Vision-Language Models (VLMs) have shown remarkable capabilities across diverse visual tasks, including image recognition, video understanding, and Visual Question Answering (VQA) when explicitly trained for these tasks.…

ObjectObject TrackingQuestion AnsweringVideo Object Tracking+3

Surgical-LVLM: Learning to Adapt Large Vision-Language Model for Grounded Visual Question Answering in Robotic Surgery

2024-03-22 · Guankun Wang, Long Bai, Wan Jun Nah, Jie Wang 외

Recent advancements in Surgical Visual Question Answering (Surgical-VQA) and related region grounding have shown great promise for robotic and medical applications, addressing the critical need for automated methods in p…

Language ModelingLanguage ModellingQuestion AnsweringVisual Grounding+2

PAL: Intelligence Augmentation using Egocentric Visual Context Detection

2021-05-22 · Mina Khan, Pattie Maes

Egocentric visual context detection can support intelligence augmentation applications. We created a wearable system, called PAL, for wearable, personalized, and privacy-preserving egocentric visual context detection. PA…

ClusteringDeep LearningFace DetectionPrivacy Preserving

INSID3: Training-Free In-Context Segmentation with DINOv3

2026-03-30 · Claudia Cuttano, Gabriele Trivigno, Christoph Reich, Daniel Cremers 외 arxiv

In-context segmentation (ICS) aims to segment arbitrary concepts, e.g., objects, parts, or personalized instances, given one annotated visual examples. Existing work relies on (i) fine-tuning vision foundation models (VF…

Personalized SegmentationSemantic correspondence