paper-with-me

Papers

DMAGaze: Gaze Estimation Based on Feature Disentanglement and Multi-Scale Attention

2025-04-15 · Haohan Chen, Hongjia Liu, Shiyong Lan, Wenwu Wang, Yixin Qiao, Yao Li, Guonan Deng

Gaze estimation, which predicts gaze direction, commonly faces the challenge of interference from complex gaze-irrelevant information in face images. In this work, we propose DMAGaze, a novel gaze estimation framework that exploits information from facial images in three aspects: gaze-relevant global features (disentangled from facial image), local eye features (extracted from cropped eye patch), and head pose estimation features, to improve overall performance. Firstly, we design a new continuous mask-based Disentangler to accurately disentangle gaze-relevant and gaze-irrelevant information in facial images by achieving the dual-branch disentanglement goal through separately reconstructing the eye and non-eye regions. Furthermore, we introduce a new cascaded attention module named Multi-Scale Global Local Attention Module (MS-GLAM). Through a customized cascaded attention structure, it effectively focuses on global and local information at multiple scales, further enhancing the information from the Disentangler. Finally, the global gaze-relevant features disentangled by the upper face branch, combined with head pose and local eye features, are passed through the detection head for high-precision gaze estimation. Our proposed DMAGaze has been extensively validated on two mainstream public datasets, achieving state-of-the-art performance.

📄 PDF Abstract BibTeX arXiv:2504.11160

Code (0)

등록된 구현이 없습니다.

Tasks

DisentanglementGaze EstimationHead Pose EstimationPose Estimation

Methods 이 논문이 사용한 방법론

Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
Attention 설명 없음
Global Local Attention Module 설명 없음

Similar Papers 제목 키워드 기반

LatentGaze: Cross-Domain Gaze Estimation through Gaze-Aware Analytic Latent Code Manipulation

2022-09-21 · Isack Lee, Jun-Seok Yun, Hee Hyeon Kim, Youngju Na 외

Although recent gaze estimation methods lay great emphasis on attentively extracting gaze-relevant features from facial or eye images, how to define features that include gaze-relevant components has been ambiguous. This…

DisentanglementGaze EstimationGenerative Adversarial Network

Investigation of Architectures and Receptive Fields for Appearance-based Gaze Estimation

2023-08-18 · Yunhan Wang, Xiangwei Shi, Shalini De Mello, Hyung Jin Chang 외

With the rapid development of deep learning technology in the past decade, appearance-based gaze estimation has attracted great attention from both computer vision and human-computer interaction research communities. Fas…

Contrastive LearningDisentanglementGaze EstimationHard Attention

Domain-Adaptive Full-Face Gaze Estimation via Novel-View-Synthesis and Feature Disentanglement

2023-05-25 · Jiawei Qin, Takuru Shimoyama, Xucong Zhang, Yusuke Sugano

Along with the recent development of deep neural networks, appearance-based gaze estimation has succeeded considerably when training and testing within the same domain. Compared to the within-domain task, the variance of…

3D ReconstructionDisentanglementDomain AdaptationGaze Estimation+2

GazeShift: Unsupervised Gaze Estimation and Dataset for VR

2026-03-08 · Gil Shapira, Ishay Goldin, Evgeny Artyomov, Donghoon Kim 외 arxiv

Gaze estimation is instrumental in modern virtual reality (VR) systems. Despite significant progress in remote-camera gaze estimation, VR gaze research remains constrained by data scarcity, particularly the lack of large…

Gaze Estimation

LISA: Language-guided Interference-aware Spatial-Frequency Attention for Driver Gaze Estimation

2026-05-17 · Jun Ma, Zhenye Yang, Ruichen Zhou, Pei Zhang 외 arxiv

Driver gaze estimation serves as a fundamental metric for evaluating driver attentiveness in modern monitoring systems. Beyond being vulnerable to sudden lighting changes and sensor noise, spatial-domain models struggle …

Gaze Estimation