paper-with-me

Papers

A Backbone Benchmarking Study on Self-supervised Learning as a Auxiliary Task with Texture-based Local Descriptors for Face Analysis

2026-03-23 · Shukesh Reddy, Abhijit Das arxiv

In this work, we benchmark with different backbones and study their impact for self-supervised learning (SSL) as an auxiliary task to blend texture-based local descriptors into feature modelling for efficient face analysis. It is established in previous work that combining a primary task and a self-supervised auxiliary task enables more robust and discriminative representation learning. We employed different shallow to deep backbones for the SSL task of Masked Auto-Encoder (MAE) as an auxiliary objective to reconstruct texture features such as local patterns alongside the primary task in local pattern SSAT (L-SSAT), ensuring robust and unbiased face analysis. To expand the benchmark, we conducted a comprehensive comparative analysis across multiple model configurations within the proposed framework. To this end, we address the three research questions: "What is the role of the backbone in performance L-SSAT?", "What type of backbone is effective for different face analysis tasks?", and "Is there any generalized backbone for effective face analysis with L-SSAT?". Towards answering these questions, we provide a detailed study and experiments. The performance evaluation demonstrates that the backbone for the proposed method is highly dependent on the downstream task, achieving average accuracies of 0.94 on FaceForensics++, 0.87 on CelebA, and 0.88 on AffectNet. For consistency of feature representation quality and generalisation capability across various face analysis paradigms, including face attribute prediction, emotion classification, and deepfake detection, there is no unified backbone.

📄 PDF Abstract BibTeX arXiv:2603.22190

Code (0)

등록된 구현이 없습니다.

Tasks

Self-Supervised LearningRepresentation LearningEmotion ClassificationDeepFake Detection

Similar Papers 제목 키워드 기반

Benchmarking Detection Transfer Learning with Vision Transformers

2021-11-22 · Yanghao Li, Saining Xie, Xinlei Chen, Piotr Dollar 외

Object detection is a central downstream task used to test if pre-trained network parameters confer benefits, such as improved accuracy or training speed. The complexity of object detection methods can make this benchmar…

Benchmarkingobject-detectionObject DetectionSelf-Supervised Learning+1

How Self-Supervised Learning Can be Used for Fine-Grained Head Pose Estimation?

2021-08-10 · Mahdi Pourmirzaei, Farzaneh Esmaili, Ebrahim Mousavi, Sasan Karamizadeh 외

The cost of head pose labeling is the main challenge of improving the fine-grained Head Pose Estimation (HPE). Although Self-Supervised Learning (SSL) can be a solution to the lack of huge amounts of labeled data, its ef…

Head Pose EstimationMulti-Task LearningPose EstimationSelf-Supervised Learning+1

Speaker-Aware Temporal Aggregation Strategies on Segment Representations for Depression Detection in Dyadic Interaction: A Benchmark Study

2026-07-03 · Anisha Pattanayak, Huang-Cheng Chou, Shrikanth Narayanan, Sudarsana Reddy Kadiri hf

Speech-based depression detection compresses features from short audio segments into one speaker-level decision, a step called temporal aggregation rarely studied on its own. Most benchmarks fix a single self-supervised …

Self-Supervised Regional and Temporal Auxiliary Tasks for Facial Action Unit Recognition

2021-07-30 · Jingwei Yan, Jingjing Wang, Qiang Li, Chunmao Wang 외

Automatic facial action unit (AU) recognition is a challenging task due to the scarcity of manual annotations. To alleviate this problem, a large amount of efforts has been dedicated to exploiting various methods which l…

Facial Action Unit DetectionOptical Flow EstimationRelation

A Closer Look at Benchmarking Self-Supervised Pre-training with Image Classification

2024-07-16 · Markus Marks, Manuel Knott, Neehar Kondapaneni, Elijah Cole 외

Self-supervised learning (SSL) is a machine learning approach where the data itself provides supervision, eliminating the need for external labels. The model is forced to learn about the data structure or context by solv…

BenchmarkingFew-Shot Learningimage-classificationImage Classification+1