paper-with-me

Papers

EgoAdapt: Enhancing Robustness in Egocentric Interactive Speaker Detection Under Missing Modalities

2026-03-18 · Xinyuan Qian, Xinjia Zhu, Alessio Brutti, Dong Liang arxiv

TTM (Talking to Me) task is a pivotal component in understanding human social interactions, aiming to determine who is engaged in conversation with the camera-wearer. Traditional models often face challenges in real-world scenarios due to missing visual data, neglecting the role of head orientation, and background noise. This study addresses these limitations by introducing EgoAdapt, an adaptive framework designed for robust egocentric "Talking to Me" speaker detection under missing modalities. Specifically, EgoAdapt incorporates three key modules: (1) a Visual Speaker Target Recognition (VSTR) module that captures head orientation as a non-verbal cue and lip movement as a verbal cue, allowing a comprehensive interpretation of both verbal and non-verbal signals to address TTM, setting it apart from tasks focused solely on detecting speaking status; (2) a Parallel Shared-weight Audio (PSA) encoder for enhanced audio feature extraction in noisy environments; and (3) a Visual Modality Missing Awareness (VMMA) module that estimates the presence or absence of each modality at each frame to adjust the system response dynamically.Comprehensive evaluations on the TTM benchmark of the Ego4D dataset demonstrate that EgoAdapt achieves a mean Average Precision (mAP) of 67.39% and an Accuracy (Acc) of 62.01%, significantly outperforming the state-of-the-art method by 4.96% in Accuracy and 1.56% in mAP.

📄 PDF Abstract BibTeX arXiv:2603.18082

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

EgoAdapt: Adaptive Multisensory Distillation and Policy Learning for Efficient Egocentric Perception

2025-06-26 · Sanjoy Chowdhury, Subrata Biswas, Sayan Nag, Tushar Nagarajan 외

Modern perception models, particularly those designed for multisensory egocentric tasks, have achieved remarkable performance but often come with substantial computational costs. These high demands pose challenges for re…

Action RecognitionActive Speaker Localization

EgoAdapt: A Multi-Scene Egocentric Adaptation Method for CVPR 2026 HD-EPIC VQA Challenge

2026-05-23 · Zhiwei Chen, Yupeng Hu, Zixu Li, Zhiheng Fu 외 arxiv

This technical report presents our solution, EgoAdapt (Egocentric Adaptation via Category, Calibration, and Consistency), to the CVPR 2026 HD-EPIC VQA challenge. HD-EPIC evaluates whether a vision-language model can reas…

EgoAdapt: A multi-stream evaluation study of adaptation to real-world egocentric user video

2023-07-11 · Matthias De Lange, Hamid Eghbalzadeh, Reuben Tan, Michael Iuzzolino 외

In egocentric action recognition a single population model is typically trained and subsequently embodied on a head-mounted device, such as an augmented reality headset. While this model remains static for new users and …

Action RecognitionContinual Learning

EgoPriMo: Egocentric Motion Generation for Interactive Humanoid Control

2026-06-07 · Haoyang Ge, Peng Ren, Yukun Shi, Cong Huang 외 arxiv

Humanoid robots require whole-body motions that adapt to scene context, task requirements, and user intent. Motion tracking reproduces specified trajectories, and humanoid vision-language-action systems provide semantic …

Egocentric Speaker Classification in Child-Adult Dyadic Interactions: From Sensing to Computational Modeling

2024-09-14 · Tiantian Feng, Anfeng Xu, Xuan Shi, Somer Bishop 외

Autism spectrum disorder (ASD) is a neurodevelopmental condition characterized by challenges in social communication, repetitive behavior, and sensory processing. One important research area in ASD is evaluating children…