paper-with-me

Papers

Audiovisual Speaker Tracking using Nonlinear Dynamical Systems with Dynamic Stream Weights

2019-03-14 · Christopher Schymura, Dorothea Kolossa

Data fusion plays an important role in many technical applications that require efficient processing of multimodal sensory observations. A prominent example is audiovisual signal processing, which has gained increasing attention in automatic speech recognition, speaker localization and related tasks. If appropriately combined with acoustic information, additional visual cues can help to improve the performance in these applications, especially under adverse acoustic conditions. A dynamic weighting of acoustic and visual streams based on instantaneous sensor reliability measures is an efficient approach to data fusion in this context. This paper presents a framework that extends the well-established theory of nonlinear dynamical systems with the notion of dynamic stream weights for an arbitrary number of sensory observations. It comprises a recursive state estimator based on the Gaussian filtering paradigm, which incorporates dynamic stream weights into a framework closely related to the extended Kalman filter. Additionally, a convex optimization approach to estimate oracle dynamic stream weights in fully observed dynamical systems utilizing a Dirichlet prior is presented. This serves as a basis for a generic parameter learning framework of dynamic stream weight estimators. The proposed system is application-independent and can be easily adapted to specific tasks and requirements. A study using audiovisual speaker tracking tasks is considered as an exemplary application in this work. An improved tracking performance of the dynamic stream weight-based estimation framework over state-of-the-art methods is demonstrated in the experiments.

📄 PDF Abstract BibTeX arXiv:1903.06031

Code (1)

rub-ksv/avtrack 공식 구현

Tasks

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

Similar Papers 제목 키워드 기반

Speech Trax: A Bottom to the Top Approach for Speaker Tracking and Indexing in an Archiving Context

2016-05-01 · LREC 2016 5 · F{\'e}licien Vallet, Jim Uro, J{\'e}r{\'e}my Andriamakaoly, Hakim Nabi 외

With the increasing amount of audiovisual and digital data deriving from televisual and radiophonic sources, professional archives such as INA, France{'}s national audiovisual institute, acknowledge a growing need for ef…

speaker-diarizationSpeaker Diarization

See, Hear, and Understand: Benchmarking Audiovisual Human Speech Understanding in Multimodal Large Language Models

2025-12-01 · Le Thien Phuc Nguyen, Zhuoran Yu, Samuel Low Yu Hang, Subin An 외 arxiv

Multimodal large language models (MLLMs) are expected to jointly interpret vision, audio, and language, yet existing video benchmarks rarely assess fine-grained reasoning about human speech. Many tasks remain visually so…

Learning Control-Oriented Dynamical Structure from Data

2023-02-06 · Spencer M. Richards, Jean-Jacques Slotine, Navid Azizan, Marco Pavone

Even for known nonlinear dynamical systems, feedback controller synthesis is a difficult problem that often requires leveraging the particular structure of the dynamics to induce a stable closed-loop system. For general …

Audiovisual speaker conversion: jointly and simultaneously transforming facial expression and acoustic characteristics

2018-10-29 · Fuming Fang, Xin Wang, Junichi Yamagishi, Isao Echizen

An audiovisual speaker conversion method is presented for simultaneously transforming the facial expressions and voice of a source speaker into those of a target speaker. Transforming the facial and acoustic features tog…

Image Reconstruction

Data Fusion for Audiovisual Speaker Localization: Extending Dynamic Stream Weights to the Spatial Domain

2021-02-23 · Julio Wissing, Benedikt Boenninghoff, Dorothea Kolossa, Tsubasa Ochiai 외

Estimating the positions of multiple speakers can be helpful for tasks like automatic speech recognition or speaker diarization. Both applications benefit from a known speaker position when, for instance, applying beamfo…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Positionspeaker-diarization+3