Spherical World-Locking for Audio-Visual Localization in Egocentric Videos
Egocentric videos provide comprehensive contexts for user and scene understanding, spanning multisensory perception to behavioral interaction. We propose Spherical World-Locking (SWL) as a general framework for egocentric scene representation, which implicitly transforms multisensory streams with respect to measurements of head orientation. Compared to conventional head-locked egocentric representations with a 2D planar field-of-view, SWL effectively offsets challenges posed by self-motion, allowing for improved spatial synchronization between input modalities. Using a set of multisensory embeddings on a worldlocked sphere, we design a unified encoder-decoder transformer architecture that preserves the spherical structure of the scene representation, without requiring expensive projections between image and world coordinate systems. We evaluate the effectiveness of the proposed framework on multiple benchmark tasks for egocentric video understanding, including audio-visual active speaker localization, auditory spherical source localization, and behavior anticipation in everyday activities.
Code (0)
등록된 구현이 없습니다.
Tasks
Active Speaker LocalizationDecoderScene UnderstandingVideo UnderstandingVisual LocalizationMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Visual Sound Localization in the Wild by Cross-Modal Interference Erasing
The task of audio-visual sound source localization has been well studied under constrained scenes, where the audio recordings are clean. However, in real-world scenarios, audios are usually contaminated by off-screen sou…
Sound Source LocalizationPano-AVQA: Grounded Audio-Visual Question Answering on 360deg Videos
360deg videos convey holistic views for the surroundings of a scene. It provides audio-visual cues beyond predetermined normal field of views and displays distinctive spatial relations on a sphere. However, previous …
Audio-visual Question AnsweringQuestion AnsweringRelationVisual Question Answering+1Pano-AVQA: Grounded Audio-Visual Question Answering on 360$^\circ$ Videos
360$^\circ$ videos convey holistic views for the surroundings of a scene. It provides audio-visual cues beyond pre-determined normal field of views and displays distinctive spatial relations on a sphere. However, previou…
Audio-visual Question AnsweringQuestion AnsweringRelationVisual Question Answering+1Spherical Vision Transformers for Audio-Visual Saliency Prediction in 360-Degree Videos
Omnidirectional videos (ODVs) are redefining viewer experiences in virtual reality (VR) by offering an unprecedented full field-of-view (FOV). This study extends the domain of saliency prediction to 360-degree environmen…
Saliency PredictionLAVA: Layered Audio-Visual Anti-tampering Watermarking for Robust Deepfake Detection and Localization
Proactive watermarking offers a promising approach for deepfake tamper detection and localization in short-form videos. However, existing methods often decouple audio and visual evidence and assume that watermark signals…
DeepFake Detection