paper-with-me

Papers

Kernel-based Sensor Fusion with Application to Audio-Visual Voice Activity Detection

2016-04-11 · David Dov, Ronen Talmon, Israel Cohen

In this paper, we address the problem of multiple view data fusion in the presence of noise and interferences. Recent studies have approached this problem using kernel methods, by relying particularly on a product of kernels constructed separately for each view. From a graph theory point of view, we analyze this fusion approach in a discrete setting. More specifically, based on a statistical model for the connectivity between data points, we propose an algorithm for the selection of the kernel bandwidth, a parameter, which, as we show, has important implications on the robustness of this fusion approach to interferences. Then, we consider the fusion of audio-visual speech signals measured by a single microphone and by a video camera pointed to the face of the speaker. Specifically, we address the task of voice activity detection, i.e., the detection of speech and non-speech segments, in the presence of structured interferences such as keyboard taps and office noise. We propose an algorithm for voice activity detection based on the audio-visual signal. Simulation results show that the proposed algorithm outperforms competing fusion and voice activity detection approaches. In addition, we demonstrate that a proper selection of the kernel bandwidth indeed leads to improved performance.

📄 PDF Abstract BibTeX arXiv:1604.02946

Code (0)

등록된 구현이 없습니다.

Tasks

Action DetectionActivity DetectionSensor Fusion

Similar Papers 제목 키워드 기반

Multi-modal Egocentric Activity Recognition using Audio-Visual Features

2018-07-02 · Mehmet Ali Arabaci, Fatih Özkan, Elif Surer, Peter Jančovič 외

Egocentric activity recognition in first-person videos has an increasing importance with a variety of applications such as lifelogging, summarization, assisted-living and activity tracking. Existing methods for this task…

Activity RecognitionEgocentric Activity RecognitionOptical Flow Estimation

Audiovisual Speaker Tracking using Nonlinear Dynamical Systems with Dynamic Stream Weights

2019-03-14 · Christopher Schymura, Dorothea Kolossa

Data fusion plays an important role in many technical applications that require efficient processing of multimodal sensory observations. A prominent example is audiovisual signal processing, which has gained increasing a…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

Pay Self-Attention to Audio-Visual Navigation

2022-10-04 · Yinfeng Yu, Lele Cao, Fuchun Sun, Xiaohong Liu 외

Audio-visual embodied navigation, as a hot research topic, aims training a robot to reach an audio target using egocentric visual (from the sensors mounted on the robot) and audio (emitted from the target) input. The aud…

Visual Navigation

STNet: Deep Audio-Visual Fusion Network for Robust Speaker Tracking

2024-10-08 · Yidi Li, Hong Liu, Bing Yang

Audio-visual speaker tracking aims to determine the location of human targets in a scene using signals captured by a multi-sensor platform, whose accuracy and robustness can be improved by multi-modal fusion methods. Rec…

Dynamic Multi-Target Fusion for Efficient Audio-Visual Navigation

2025-09-23 · Yinfeng Yu, Hailong Zhang, Meiling Zhu arxiv

Audiovisual embodied navigation enables robots to locate audio sources by dynamically integrating visual observations from onboard sensors with the auditory signals emitted by the target. The core challenge lies in effec…

Visual Navigation