paper-with-me

홈 › Papers

Not All Birds Look The Same: Identity-Preserving Generation For Birds

2025-12-04 · Aaron Sun, Oindrila Saha, Subhransu Maji arxiv

Since the advent of controllable image generation, increasingly rich modes of control have enabled greater customization and accessibility for everyday users. Zero-shot, identity-preserving models such as Insert Anything and OminiControl now support applications like virtual try-on without requiring additional fine-tuning. While these models may be fitting for humans and rigid everyday objects, they still have limitations for non-rigid or fine-grained categories. These domains often lack accessible, high-quality data -- especially videos or multi-view observations of the same subject -- making them difficult both to evaluate and to improve upon. Yet, such domains are essential for moving beyond content creation toward applications that demand accuracy and fine detail. Birds are an excellent domain for this task: they exhibit high diversity, require fine-grained cues for identification, and come in a wide variety of poses. We introduce the NABirds Look-Alikes (NABLA) dataset, consisting of 4,759 expert-curated image pairs. Together with 1,073 pairs collected from multi-image observations on iNaturalist and a small set of videos, this forms a benchmark for evaluating identity-preserving generation of birds. We show that state-of-the-art baselines fail to maintain identity on this dataset, and we demonstrate that training on images grouped by species, age, and sex -- used as a proxy for identity -- substantially improves performance on both seen and unseen species.

📄 PDF Abstract BibTeX arXiv:2512.04485

Code (0)

등록된 구현이 없습니다.

Tasks

Image GenerationVirtual Try-on

Similar Papers 제목 키워드 기반

T-Person-GAN: Text-to-Person Image Generation with Identity-Consistency and Manifold Mix-Up

2022-08-18 · Deyin Liu, Lin Yuanbo Wu, Bo Li, ZongYuan Ge

In this paper, we present an end-to-end approach to generate high-resolution person images conditioned on texts only. State-of-the-art text-to-image generation models are mainly designed for center-object generation, e.g…

Image GenerationText to Image GenerationText-to-Image Generation

Lookahead Anchoring: Preserving Character Identity in Audio-Driven Human Animation

2025-10-27 · Junyoung Seo, Rodrigo Mira, Alexandros Haliassos, Stella Bounareli 외 arxiv

Audio-driven human animation models often suffer from identity drift during temporal autoregressive generation, where characters gradually lose their identity over time. One solution is to generate keyframes as intermedi…

Two Birds with One Stone: Transforming and Generating Facial Images with Iterative GAN

2017-11-16 · Dan Ma, Bin Liu, Zhao Kang, Jiayu Zhou 외

Generating high fidelity identity-preserving faces with different facial attributes has a wide range of applications. Although a number of generative models have been developed to tackle this problem, there is still much…

Image Generation

MotionCharacter: Identity-Preserving and Motion Controllable Human Video Generation

2024-11-27 · Haopeng Fang, Di Qiu, Binjie Mao, Pengfei Yan 외

Recent advancements in personalized Text-to-Video (T2V) generation highlight the importance of integrating character-specific identities and actions. However, previous T2V models struggle with identity consistency and co…

AttributeVideo Generation

Identity and Attribute Preserving Thumbnail Upscaling

2021-05-30 · Noam Gat, Sagie Benaim, Lior Wolf

We consider the task of upscaling a low resolution thumbnail image of a person, to a higher resolution image, which preserves the person's identity and other attributes. Since the thumbnail image is of low resolution, ma…

Attribute