paper-with-me

Papers

How Self-Supervised Learning Can be Used for Fine-Grained Head Pose Estimation?

2021-08-10 · Mahdi Pourmirzaei, Farzaneh Esmaili, Ebrahim Mousavi, Sasan Karamizadeh, Seyedehsamaneh Shojaeilangari

The cost of head pose labeling is the main challenge of improving the fine-grained Head Pose Estimation (HPE). Although Self-Supervised Learning (SSL) can be a solution to the lack of huge amounts of labeled data, its efficacy for fine-grained HPE is not yet fully explored. This study aims to assess the usage of SSL in fine-grained HPE based on two scenarios: (1) using SSL for weights pre-training procedure, and (2) leveraging auxiliary SSL losses besides HPE. We design a Hybrid Multi-Task Learning (HMTL) architecture based on the ResNet50 backbone in which both strategies are applied. Our experimental results reveal that the combination of both scenarios is the best for HPE. Together, the average error rate is reduced up to 23.1% for AFLW2000 and 14.2% for BIWI benchmark compared to the baseline. Moreover, it is found that some SSL methods are more suitable for transfer learning, while others may be effective when they are considered as auxiliary tasks incorporated into supervised learning. Finally, it is shown that by using the proposed HMTL architecture, the average error is reduced with different types of initial weights: random, ImageNet and SSL pre-trained weights.

📄 PDF Abstract BibTeX arXiv:2108.04893

Code (0)

등록된 구현이 없습니다.

Tasks

Head Pose EstimationMulti-Task LearningPose EstimationSelf-Supervised LearningTransfer Learning

Methods 이 논문이 사용한 방법론

Jigsaw Jigsaw is a self-supervision approach that relies on jigsaw-like puzzles as the pretext task in order to learn image representations.

Similar Papers 제목 키워드 기반

Using Self-Supervised Auxiliary Tasks to Improve Fine-Grained Facial Representation

2021-05-13 · Mahdi Pourmirzaei, Gholam Ali Montazer, Farzaneh Esmaili

In this paper, at first, the impact of ImageNet pre-training on fine-grained Facial Emotion Recognition (FER) is investigated which shows that when enough augmentations on images are applied, training from scratch provid…

Emotion RecognitionFacial Emotion RecognitionFacial Expression Recognition (FER)Head Pose Estimation+3

Enhancing Self-Supervised Talking Head Forgery Detection via a Training-Free Dual-System Framework

2026-05-05 · Ke Liu, Jiwei Wei, Shuchang Zhou, Yutong Xiao 외 arxiv

Supervised talking head forgery detection faces severe generalization challenges due to the continuous evolution of generators. By reducing reliance on generator-specific forgery patterns, self-supervised detectors offer…

Towards Fine-grained Visual Representations by Combining Contrastive Learning with Image Reconstruction and Attention-weighted Pooling

2021-04-09 · Jonas Dippel, Steffen Vogler, Johannes Höhne

This paper presents Contrastive Reconstruction, ConRec - a self-supervised learning algorithm that obtains image representations by jointly optimizing a contrastive and a self-reconstruction loss. We showcase that state-…

Contrastive LearningDecoderImage ReconstructionSelf-Supervised Learning

Progressive Multi-Scale Self-Supervised Learning for Speech Recognition

2022-12-07 · Genshun Wan, Tan Liu, Hang Chen, Jia Pan 외

Self-supervised learning (SSL) models have achieved considerable improvements in automatic speech recognition (ASR). In addition, ASR performance could be further improved if the model is dedicated to audio content infor…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Self-Supervised Learningspeech-recognition+1

Fine-Grained Self-Supervised Learning with Jigsaw Puzzles for Medical Image Classification

2023-08-10 · Wongi Park, Jongbin Ryu

Classifying fine-grained lesions is challenging due to minor and subtle differences in medical images. This is because learning features of fine-grained lesions with highly minor differences is very difficult in training…

image-classificationImage ClassificationMedical Image ClassificationSelf-Supervised Learning