Active Speaker Localization
1개 벤치마크 · 논문 5편 · 이 태스크의 논문 보기 →
Benchmarks
EasyCom
Papers
EgoAdapt: Adaptive Multisensory Distillation and Policy Learning for Efficient Egocentric Perception
Modern perception models, particularly those designed for multisensory egocentric tasks, have achieved remarkable performance but often come with substantial computational costs. These high demands pose challenges for re…
Action RecognitionActive Speaker LocalizationSpherical World-Locking for Audio-Visual Localization in Egocentric Videos
Egocentric videos provide comprehensive contexts for user and scene understanding, spanning multisensory perception to behavioral interaction. We propose Spherical World-Locking (SWL) as a general framework for egocentri…
Active Speaker LocalizationDecoderScene UnderstandingVideo Understanding+1Audio visual character profiles for detecting background characters in entertainment media
An essential goal of computational media intelligence is to support understanding how media stories -- be it news, commercial or entertainment media -- represent and reflect society and these portrayals are perceived. Pe…
Active Speaker LocalizationFace VerificationEgocentric Deep Multi-Channel Audio-Visual Active Speaker Localization
Augmented reality devices have the potential to enhance human perception and enable other assistive functionalities in complex conversational environments. Effectively capturing the audio-visual context necessary for und…
Action DetectionActive Speaker DetectionActive Speaker LocalizationActivity Detection+1Cross modal video representations for weakly supervised active speaker localization
An objective understanding of media depictions, such as inclusive portrayals of how much someone is heard and seen on screen such as in film and television, requires the machines to discern automatically who, when, how, …
Action DetectionActive Speaker LocalizationActivity DetectionEvent Detection