paper-with-me

Papers

Omni-ID: Holistic Identity Representation Designed for Generative Tasks

2024-12-12 · CVPR 2025 1 · Guocheng Qian, Kuan-Chieh Wang, Or Patashnik, Negin Heravi, Daniil Ostashev, Sergey Tulyakov, Daniel Cohen-Or, Kfir Aberman

We introduce Omni-ID, a novel facial representation designed specifically for generative tasks. Omni-ID encodes holistic information about an individual's appearance across diverse expressions and poses within a fixed-size representation. It consolidates information from a varied number of unstructured input images into a structured representation, where each entry represents certain global or local identity features. Our approach uses a few-to-many identity reconstruction training paradigm, where a limited set of input images is used to reconstruct multiple target images of the same individual in various poses and expressions. A multi-decoder framework is further employed to leverage the complementary strengths of diverse decoders during training. Unlike conventional representations, such as CLIP and ArcFace, which are typically learned through discriminative or contrastive objectives, Omni-ID is optimized with a generative objective, resulting in a more comprehensive and nuanced identity capture for generative tasks. Trained on our MFHQ dataset -- a multi-view facial image collection, Omni-ID demonstrates substantial improvements over conventional representations across various generative tasks.

📄 PDF Abstract BibTeX arXiv:2412.09694

Code (0)

등록된 구현이 없습니다.

Tasks

Decoder

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically
CLIP Contrastive Language-Image Pre-training (CLIP), consisting of a simplified version of ConVIRT trained from scratch, is an efficient method of image representation learning…
ArcFace ArcFace, or Additive Angular Margin Loss, is a loss function used in face recognition tasks. The softmax is traditionally used…

Similar Papers 제목 키워드 기반

Omni-Attribute: Open-vocabulary Attribute Encoder for Visual Concept Personalization

2025-12-11 · Tsai-Shien Chen, Aliaksandr Siarohin, Gordon Guocheng Qian, Kuan-Chieh Jackson Wang 외 arxiv

Visual concept personalization aims to transfer only specific image attributes, such as identity, expression, lighting, and style, into unseen contexts. However, existing methods rely on holistic embeddings from general-…

OmniPerson: Unified Identity-Preserving Pedestrian Generation

2025-12-02 · Changxiao Ma, Chao Yuan, Xincheng Shi, Yuzhuo Ma 외 arxiv

Person re-identification (ReID) suffers from a lack of large-scale high-quality training data due to challenges in data privacy and annotation costs. While previous approaches have explored pedestrian generation for data…

Person Re-IdentificationImage Super-ResolutionData AugmentationVideo Generation

Kling-Omni Technical Report

2025-12-18 · Kling Team, Jialu Chen, Yuanzheng Ci, Xiangyu Du 외 arxiv

We present Kling-Omni, a generalist generative framework designed to synthesize high-fidelity videos directly from multimodal visual language inputs. Adopting an end-to-end perspective, Kling-Omni bridges the functional …

Instruction FollowingVideo Generation

Omni-I2C: A Holistic Benchmark for High-Fidelity Image-to-Code Generation

2026-03-18 · Jiawei Zhou, Chi Zhang, Xiang Feng, Qiming Zhang 외 arxiv

We present Omni-I2C, a comprehensive benchmark designed to evaluate the capability of Large Multimodal Models (LMMs) in converting complex, structured digital graphics into executable code. We argue that this task repres…

Code Generation

OmniPart: Part-Aware 3D Generation with Semantic Decoupling and Structural Cohesion

2025-07-08 · Yunhan Yang, Yufan Zhou, Yuan-Chen Guo, Zi-Xin Zou 외

The creation of 3D assets with explicit, editable part structures is crucial for advancing interactive applications, yet most generative methods produce only monolithic shapes, limiting their utility. We introduce OmniPa…

3D Generation