paper-with-me

Papers

Video Joint-Embedding Predictive Architectures for Facial Expression Recognition

2026-01-14 · Lennart Eing, Cristina Luna-Jiménez, Silvan Mertes, Elisabeth André arxiv

This paper introduces a novel application of Video Joint-Embedding Predictive Architectures (V-JEPAs) for Facial Expression Recognition (FER). Departing from conventional pre-training methods for video understanding that rely on pixel-level reconstructions, V-JEPAs learn by predicting embeddings of masked regions from the embeddings of unmasked regions. This enables the trained encoder to not capture irrelevant information about a given video like the color of a region of pixels in the background. Using a pre-trained V-JEPA video encoder, we train shallow classifiers using the RAVDESS and CREMA-D datasets, achieving state-of-the-art performance on RAVDESS and outperforming all other vision-based methods on CREMA-D (+1.48 WAR). Furthermore, cross-dataset evaluations reveal strong generalization capabilities, demonstrating the potential of purely embedding-based pre-training approaches to advance FER. We release our code at https://github.com/lennarteingunia/vjepa-for-fer.

📄 PDF Abstract BibTeX arXiv:2601.09524

Code (0)

등록된 구현이 없습니다.

Tasks

Facial Expression Recognition

Similar Papers 제목 키워드 기반

MJEPA: A Simple and Scalable Joint-Embedding Predictive Architecture for Audio-Visual Learning

2026-06-23 · Revant Teotia, Adrien Bardes, Michael Rabbat, Sumit Chopra 외 arxiv

Self-supervised learning from large-scale video data has emerged as a dominant paradigm for visual representation learning. Since audio and visual streams naturally co-occur in video data, extending this success to joint…

Self-Supervised LearningRepresentation Learning

Video Representation Learning with Joint-Embedding Predictive Architectures

2024-12-14 · Katrina Drozdov, Ravid Shwartz-Ziv, Yann Lecun

Video representation learning is an increasingly important topic in machine learning research. We present Video JEPA with Variance-Covariance Regularization (VJ-VCR): a joint-embedding predictive architecture for self-su…

Representation Learning

Rhythm-Structured Predictive Learning for Remote Photoplethysmography

2026-06-30 · Ba-Thinh Nguyen, Huu-Dung Nguyen, Thi-Duyen Ngo, Thanh-Ha Le arxiv

Remote photoplethysmography (rPPG) estimates physiological signals from facial videos by analyzing subtle pulse induced skin color variations. Despite recent progress, existing self-supervised rPPG methods mainly reconst…

Representation Learning

Denoising with a Joint-Embedding Predictive Architecture

2024-10-02 · Dengsheng Chen, Jie Hu, Xiaoming Wei, Enhua Wu

Joint-embedding predictive architectures (JEPAs) have shown substantial promise in self-supervised representation learning, yet their application in generative modeling remains underexplored. Conversely, diffusion models…

DenoisingImage GenerationRepresentation Learning

Emotion-Conditioned Short-Horizon Human Pose Forecasting with a Lightweight Predictive World Model

2026-04-26 · Jingni Huang, Peter Bloodsworth arxiv

Short-term human pose prediction plays a crucial role in interactive systems, assistive robots, and emotion-aware human-computer interaction[1-3]. While current trajectory prediction models primarily rely on geometric mo…

Human Pose ForecastingTrajectory PredictionPose Prediction