paper-with-me

Papers

Spherical World-Locking for Audio-Visual Localization in Egocentric Videos

2024-08-09 · Heeseung Yun, Ruohan Gao, Ishwarya Ananthabhotla, Anurag Kumar, Jacob Donley, Chao Li, Gunhee Kim, Vamsi Krishna Ithapu, Calvin Murdock

Egocentric videos provide comprehensive contexts for user and scene understanding, spanning multisensory perception to behavioral interaction. We propose Spherical World-Locking (SWL) as a general framework for egocentric scene representation, which implicitly transforms multisensory streams with respect to measurements of head orientation. Compared to conventional head-locked egocentric representations with a 2D planar field-of-view, SWL effectively offsets challenges posed by self-motion, allowing for improved spatial synchronization between input modalities. Using a set of multisensory embeddings on a worldlocked sphere, we design a unified encoder-decoder transformer architecture that preserves the spherical structure of the scene representation, without requiring expensive projections between image and world coordinate systems. We evaluate the effectiveness of the proposed framework on multiple benchmark tasks for egocentric video understanding, including audio-visual active speaker localization, auditory spherical source localization, and behavior anticipation in everyday activities.

📄 PDF Abstract BibTeX arXiv:2408.05364

Code (0)

등록된 구현이 없습니다.

Tasks

Active Speaker LocalizationDecoderScene UnderstandingVideo UnderstandingVisual Localization

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically

Similar Papers 제목 키워드 기반

Visual Sound Localization in the Wild by Cross-Modal Interference Erasing

2022-02-13 · Xian Liu, Rui Qian, Hang Zhou, Di Hu 외

The task of audio-visual sound source localization has been well studied under constrained scenes, where the audio recordings are clean. However, in real-world scenarios, audios are usually contaminated by off-screen sou…

Sound Source Localization

Pano-AVQA: Grounded Audio-Visual Question Answering on 360deg Videos

2021-01-01 · ICCV 2021 10 · Heeseung Yun, Youngjae Yu, Wonsuk Yang, Kangil Lee 외

360deg videos convey holistic views for the surroundings of a scene. It provides audio-visual cues beyond predetermined normal field of views and displays distinctive spatial relations on a sphere. However, previous …

Audio-visual Question AnsweringQuestion AnsweringRelationVisual Question Answering+1

Pano-AVQA: Grounded Audio-Visual Question Answering on 360$^\circ$ Videos

2021-10-11 · Heeseung Yun, Youngjae Yu, Wonsuk Yang, Kangil Lee 외

360$^\circ$ videos convey holistic views for the surroundings of a scene. It provides audio-visual cues beyond pre-determined normal field of views and displays distinctive spatial relations on a sphere. However, previou…

Audio-visual Question AnsweringQuestion AnsweringRelationVisual Question Answering+1

Spherical Vision Transformers for Audio-Visual Saliency Prediction in 360-Degree Videos

2025-08-27 · Mert Cokelek, Halit Ozsoy, Nevrez Imamoglu, Cagri Ozcinar 외 arxiv

Omnidirectional videos (ODVs) are redefining viewer experiences in virtual reality (VR) by offering an unprecedented full field-of-view (FOV). This study extends the domain of saliency prediction to 360-degree environmen…

Saliency Prediction

LAVA: Layered Audio-Visual Anti-tampering Watermarking for Robust Deepfake Detection and Localization

2026-04-27 · Bokang Zeng, Zheng Gao, Xiaoyu Li, Xiaoyan Feng 외 arxiv

Proactive watermarking offers a promising approach for deepfake tamper detection and localization in short-form videos. However, existing methods often decouple audio and visual evidence and assume that watermark signals…

DeepFake Detection