paper-with-me

홈 › Papers

GazeMoE: Perception of Gaze Target with Mixture-of-Experts

2026-03-06 · Zhuangzhuang Dai, Zhongxi Lu, Vincent G. Zakka, Luis J. Manso, Jose M Alcaraz Calero, Chen Li arxiv

Estimating human gaze target from visible images is a critical task for robots to understand human attention, yet the development of generalizable neural architectures and training paradigms remains challenging. While recent advances in pre-trained vision foundation models offer promising avenues for locating gaze targets, the integration of multi-modal cues -- including eyes, head poses, gestures, and contextual features -- demands adaptive and efficient decoding mechanisms. Inspired by Mixture-of-Experts (MoE) for adaptive domain expertise in large vision-language models, we propose GazeMoE, a novel end-to-end framework that selectively leverages gaze-target-related cues from a frozen foundation model through MoE modules. To address class imbalance in gaze target classification (in-frame vs. out-of-frame) and enhance robustness, GazeMoE incorporates a class-balancing auxiliary loss alongside strategic data augmentations, including region-specific cropping and photometric transformations. Extensive experiments on benchmark datasets demonstrate that our GazeMoE achieves state-of-the-art performance, outperforming existing methods on challenging gaze estimation tasks. The code and pre-trained models are released at https://huggingface.co/zdai257/GazeMoE

📄 PDF Abstract BibTeX arXiv:2603.06256

Code (0)

등록된 구현이 없습니다.

Tasks

Gaze Estimation

Similar Papers 제목 키워드 기반

Domain-Expert-Guided Hybrid Mixture-of-Experts for Medical AI: Integrating Data-Driven Learning with Clinical Priors

2026-01-25 · Jinchen Gu, Nan Zhao, Lei Qiu, Lu Zhang arxiv

Mixture-of-Experts (MoE) models increase representational capacity with modest computational cost, but their effectiveness in specialized domains such as medicine is limited by small datasets. In contrast, clinical pract…

GazeFormer-MoE: Context-Aware Gaze Estimation via CLIP and MoE Transformer

2026-01-18 · Xinyuan Zhao, Xianrui Chen, Ahmad Chaddad arxiv

We present a semantics modulated, multi scale Transformer for 3D gaze estimation. Our model conditions CLIP global features with learnable prototype banks (illumination, head pose, background, direction), fuses these pro…

Gaze Estimation

Data-centric Design of Learning-based Surgical Gaze Perception Models in Multi-Task Simulation

2026-02-09 · Yizhou Li, Shuyuan Yang, Jiaji Su, Zonghe Chua arxiv

In robot-assisted minimally invasive surgery (RMIS), reduced haptic feedback and depth cues increase reliance on expert visual perception, motivating gaze-guided training and learning-based surgical perception models. Ho…

Optimize Surgical Triplet Recognition: A Knowledge-Driven Mixture-of-Experts Solution

2026-08-24 · Yiyi Zhang, Yuchen Yuan, Ying Zheng, Jialun Pei 외 arxiv

Surgical action triplet recognition constitutes a critical task in context-aware robot-assisted surgery, facilitating automatic surgical action perception by identifying instrument, verb, target, and their association. H…

Action Triplet Recognition

Towards Pixel-Level Prediction for Gaze Following: Benchmark and Approach

2024-11-30 · Feiyang Liu, Dan Guo, Jingyuan Xu, Zihao He 외

Following the gaze of other people and analyzing the target they are looking at can help us understand what they are thinking, and doing, and predict the actions that may follow. Existing methods for gaze following strug…

Segmentation