paper-with-me

Papers

Event-based Facial Keypoint Alignment via Cross-Modal Fusion Attention and Self-Supervised Multi-Event Representation Learning

2025-09-29 · Donghwa Kang, Junho Kim, Dongwoo Kang arxiv

Event cameras offer unique advantages for facial keypoint alignment under challenging conditions, such as low light and rapid motion, due to their high temporal resolution and robustness to varying illumination. However, existing RGB facial keypoint alignment methods do not perform well on event data, and training solely on event data often leads to suboptimal performance because of its limited spatial information. Moreover, the lack of comprehensive labeled event datasets further hinders progress in this area. To address these issues, we propose a novel framework based on cross-modal fusion attention (CMFA) and self-supervised multi-event representation learning (SSMER) for event-based facial keypoint alignment. Our framework employs CMFA to integrate corresponding RGB data, guiding the model to extract robust facial features from event input images. In parallel, SSMER enables effective feature learning from unlabeled event data, overcoming spatial limitations. Extensive experiments on our real-event E-SIE dataset and a synthetic-event version of the public WFLW-V benchmark show that our approach consistently surpasses state-of-the-art methods across multiple evaluation metrics.

📄 PDF Abstract BibTeX arXiv:2509.24968

Code (0)

등록된 구현이 없습니다.

Tasks

Representation Learning

Similar Papers 제목 키워드 기반

Retinal IPA: Iterative KeyPoints Alignment for Multimodal Retinal Imaging

2024-07-25 · Jiacheng Wang, Hao Li, Dewei Hu, Rui Xu 외

We propose a novel framework for retinal feature point alignment, designed for learning cross-modality features to enhance matching and registration across multi-modality retinal images. Our model draws on the success of…

Segmentation

Facial Keypoint Sequence Generation from Audio

2020-11-02 · Prateek Manocha, Prithwijit Guha

Whenever we speak, our voice is accompanied by facial movements and expressions. Several recent works have shown the synthesis of highly photo-realistic videos of talking faces, but they either require a source video to …

A cross-modal network for facial expression recognition

2026-05-06 · Chunwei Tian, Jingyuan Xie, Qi Zhang, Chao Li 외 arxiv

Deep neural networks enriched with structural information have been widely employed for facial expression recognition tasks. However, these methods often depend on hierarchical information rather than face property to fi…

Facial Expression RecognitionFace Alignment

SuperEvent: Cross-Modal Learning of Event-based Keypoint Detection

2025-03-31 · Yannick Burkhardt, Simon Schaefer, Stefan Leutenegger

Event-based keypoint detection and matching holds significant potential, enabling the integration of event sensors into highly optimized Visual SLAM systems developed for frame cameras over decades of research. Unfortuna…

Keypoint Detection

EI-Nexus: Towards Unmediated and Flexible Inter-Modality Local Feature Extraction and Matching for Event-Image Data

2024-10-29 · Zhonghua Yi, Hao Shi, Qi Jiang, Kailun Yang 외

Event cameras, with high temporal resolution and high dynamic range, have limited research on the inter-modality local feature extraction and matching of event-image data. We propose EI-Nexus, an unmediated and flexible …

Pose Estimation