paper-with-me

Papers

A proto-object based audiovisual saliency map

2020-03-15 · Sudarshan Ramenahalli

Natural environment and our interaction with it is essentially multisensory, where we may deploy visual, tactile and/or auditory senses to perceive, learn and interact with our environment. Our objective in this study is to develop a scene analysis algorithm using multisensory information, specifically vision and audio. We develop a proto-object based audiovisual saliency map (AVSM) for the analysis of dynamic natural scenes. A specialized audiovisual camera with $360 \degree$ Field of View, capable of locating sound direction, is used to collect spatiotemporally aligned audiovisual data. We demonstrate that the performance of proto-object based audiovisual saliency map in detecting and localizing salient objects/events is in agreement with human judgment. In addition, the proto-object based AVSM that we compute as a linear combination of visual and auditory feature conspicuity maps captures a higher number of valid salient events compared to unisensory saliency maps. Such an algorithm can be useful in surveillance, robotic navigation, video compression and related applications.

📄 PDF Abstract BibTeX arXiv:2003.06779

Code (0)

등록된 구현이 없습니다.

Tasks

ObjectvalidVideo Compression

Similar Papers 제목 키워드 기반

STAViS: Spatio-Temporal AudioVisual Saliency Network

2020-01-09 · CVPR 2020 6 · Antigoni Tsiami, Petros Koutras, Petros Maragos

We introduce STAViS, a spatio-temporal audiovisual saliency network that combines spatio-temporal visual and auditory information in order to efficiently address the problem of saliency estimation in videos. Our approach…

Saliency Prediction

Audiovisual Saliency Prediction in Uncategorized Video Sequences based on Audio-Video Correlation

2021-01-07 · Maryam Qamar Butt, Anis Ur Rahman

Substantial research has been done in saliency modeling to develop intelligent machines that can perceive and interpret their surroundings. But existing models treat videos as merely image sequences excluding any audio i…

Saliency Prediction

Investigations on End-to-End Audiovisual Fusion

2018-04-30 · Michael Wand, Ngoc Thang Vu, Juergen Schmidhuber

Audiovisual speech recognition (AVSR) is a method to alleviate the adverse effect of noise in the acoustic signal. Leveraging recent developments in deep neural network-based speech recognition, we present an AVSR neural…

speech-recognitionSpeech Recognition

An audiovisual political speech analysis incorporating eye-tracking and perception data

2012-05-01 · LREC 2012 5 · Stefan Scherer, Georg Layher, John Kane, Heiko Neumann 외

We investigate the influence of audiovisual features on the perception of speaking style and performance of politicians, utilizing a large publicly available dataset of German parliament recordings. We conduct a human pe…

Persuasiveness

Audiovisual Database with 360 Video and Higher-Order Ambisonics Audio for Perception, Cognition, Behavior, and QoE Evaluation Research

2022-12-27 · Thomas Robotham, Ashutosh Singla, Olli S. Rummukainen, Alexander Raake 외

Research into multi-modal perception, human cognition, behavior, and attention can benefit from high-fidelity content that may recreate real-life-like scenes when rendered on head-mounted displays. Moreover, aspects of a…