paper-with-me

Papers

ProS: Facial Omni-Representation Learning via Prototype-based Self-Distillation

2023-11-03 · Xing Di, Yiyu Zheng, Xiaoming Liu, Yu Cheng

This paper presents a novel approach, called Prototype-based Self-Distillation (ProS), for unsupervised face representation learning. The existing supervised methods heavily rely on a large amount of annotated training facial data, which poses challenges in terms of data collection and privacy concerns. To address these issues, we propose ProS, which leverages a vast collection of unlabeled face images to learn a comprehensive facial omni-representation. In particular, ProS consists of two vision-transformers (teacher and student models) that are trained with different augmented images (cropping, blurring, coloring, etc.). Besides, we build a face-aware retrieval system along with augmentations to obtain the curated images comprising predominantly facial areas. To enhance the discrimination of learned features, we introduce a prototype-based matching loss that aligns the similarity distributions between features (teacher or student) and a set of learnable prototypes. After pre-training, the teacher vision transformer serves as a backbone for downstream tasks, including attribute estimation, expression recognition, and landmark alignment, achieved through simple fine-tuning with additional layers. Extensive experiments demonstrate that our method achieves state-of-the-art performance on various tasks, both in full and few-shot settings. Furthermore, we investigate pre-training with synthetic face images, and ProS exhibits promising performance in this scenario as well.

📄 PDF Abstract BibTeX arXiv:2311.01929

Code (0)

등록된 구현이 없습니다.

Tasks

AttributeRepresentation Learning

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically
Multi-Head Attention 설명 없음
Attention 설명 없음
Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
Residual Connection 설명 없음
Dense Connections Dense Connections, or Fully Connected Connections, are a type of layer in a deep neural network that use a linear operation where every input is connected to every output…
Layer Normalization Unlike batch normalization, Layer Normalization directly estimates the normalization statistics from the summed inputs…

Similar Papers 제목 키워드 기반

Real-time Multi-view Omnidirectional Depth Estimation System for Robots and Autonomous Driving on Real Scenes

2024-09-12 · Ming Li, Xiong Yang, Chaofan Wu, Jiaheng Li 외

Omnidirectional Depth Estimation has broad application prospects in fields such as robotic navigation and autonomous driving. In this paper, we propose a robotic prototype system and corresponding algorithm designed to v…

Autonomous DrivingDepth EstimationEdge-computing

Ex-Omni: Enabling 3D Facial Animation Generation for Omni-modal Large Language Models

2026-02-06 · Haoyu Zhang, Zhipeng Li, Yiwen Guo, Tianshu Yu arxiv

Omni-modal large language models (OLLMs) aim to unify multimodal understanding and generation, yet extending them to jointly produce speech and 3D facial animation remains largely unexplored despite its importance for na…

A Generalist FaceX via Learning Unified Facial Representation

2023-12-31 · Yue Han, Jiangning Zhang, Junwei Zhu, Xiangtai Li 외

This work presents FaceX framework, a novel facial generalist model capable of handling diverse facial tasks simultaneously. To achieve this goal, we initially formulate a unified facial representation for a broad spectr…

Facial Editing

Selfie Taking with Facial Expression Recognition Using Omni-directional Camera

2024-05-25 · Kazutaka Kiuchi, Shimpei Imamura, Norihiko Kawai

Recent studies have shown that visually impaired people have desires to take selfies in the same way as sighted people do to record their photos and share them with others. Although support applications using sound and v…

AllFace DetectionFacial Expression Recognition

Omni-ID: Holistic Identity Representation Designed for Generative Tasks

2024-12-12 · CVPR 2025 1 · Guocheng Qian, Kuan-Chieh Wang, Or Patashnik, Negin Heravi 외

We introduce Omni-ID, a novel facial representation designed specifically for generative tasks. Omni-ID encodes holistic information about an individual's appearance across diverse expressions and poses within a fixed-si…

Decoder