paper-with-me

Papers Active Speaker Localization

“Active Speaker Localization” 태그가 달린 논문 5편 · 필터 해제

EgoAdapt: Adaptive Multisensory Distillation and Policy Learning for Efficient Egocentric Perception

2025-06-26 · Sanjoy Chowdhury, Subrata Biswas, Sayan Nag, Tushar Nagarajan 외

Modern perception models, particularly those designed for multisensory egocentric tasks, have achieved remarkable performance but often come with substantial computational costs. These high demands pose challenges for re…

Action RecognitionActive Speaker Localization

Spherical World-Locking for Audio-Visual Localization in Egocentric Videos

2024-08-09 · Heeseung Yun, Ruohan Gao, Ishwarya Ananthabhotla, Anurag Kumar 외

Egocentric videos provide comprehensive contexts for user and scene understanding, spanning multisensory perception to behavioral interaction. We propose Spherical World-Locking (SWL) as a general framework for egocentri…

Active Speaker LocalizationDecoderScene UnderstandingVideo Understanding+1

Audio visual character profiles for detecting background characters in entertainment media

2022-03-21 · Rahul Sharma, Shrikanth Narayanan

An essential goal of computational media intelligence is to support understanding how media stories -- be it news, commercial or entertainment media -- represent and reflect society and these portrayals are perceived. Pe…

Active Speaker LocalizationFace Verification

Egocentric Deep Multi-Channel Audio-Visual Active Speaker Localization

2022-01-06 · CVPR 2022 1 · Hao Jiang, Calvin Murdock, Vamsi Krishna Ithapu

Augmented reality devices have the potential to enhance human perception and enable other assistive functionalities in complex conversational environments. Effectively capturing the audio-visual context necessary for und…

Action DetectionActive Speaker DetectionActive Speaker LocalizationActivity Detection+1

Cross modal video representations for weakly supervised active speaker localization

2020-03-09 · Rahul Sharma, Krishna Somandepalli, Shrikanth Narayanan

An objective understanding of media depictions, such as inclusive portrayals of how much someone is heard and seen on screen such as in film and television, requires the machines to discern automatically who, when, how, …

Action DetectionActive Speaker LocalizationActivity DetectionEvent Detection
1–5 / 5