paper-with-me

Papers

Audio-Driven Reinforcement Learning for Head-Orientation in Naturalistic Environments

2024-09-16 · Wessel Ledder, Yuzhen Qin, Kiki van der Heijden

Although deep reinforcement learning (DRL) approaches in audio signal processing have seen substantial progress in recent years, audio-driven DRL for tasks such as navigation, gaze control and head-orientation control in the context of human-robot interaction have received little attention. Here, we propose an audio-driven DRL framework in which we utilise deep Q-learning to develop an autonomous agent that orients towards a talker in the acoustic environment based on stereo speech recordings. Our results show that the agent learned to perform the task at a near perfect level when trained on speech segments in anechoic environments (that is, without reverberation). The presence of reverberation in naturalistic acoustic environments affected the agent's performance, although the agent still substantially outperformed a baseline, randomly acting agent. Finally, we quantified the degree of generalization of the proposed DRL approach across naturalistic acoustic environments. Our experiments revealed that policies learned by agents trained on medium or high reverb environments generalized to low reverb environments, but policies learned by agents trained on anechoic or low reverb environments did not generalize to medium or high reverb environments. Taken together, this study demonstrates the potential of audio-driven DRL for tasks such as head-orientation control and highlights the need for training strategies that enable robust generalization across environments for real-world audio-driven DRL applications.

📄 PDF Abstract BibTeX arXiv:2409.10048

Code (1)

humanandmachinehearing/audiodriven_drl_for_headorientationcontrol 공식 구현 pytorch

Tasks

Audio Signal ProcessingDeep Reinforcement LearningQ-Learningreinforcement-learningReinforcement Learning

Methods 이 논문이 사용한 방법론

Q-Learning Q-Learning is an off-policy temporal difference control algorithm: $$Q\left(S\_{t}, A\_{t}\right) \leftarrow Q\left(S\_{t}, A\_{t}\right) + \alpha\left[R_{t+1} +…

Similar Papers 제목 키워드 기반

Learning Frame-Wise Emotion Intensity for Audio-Driven Talking-Head Generation

2024-09-29 · Jingyi Xu, Hieu Le, Zhixin Shu, Yang Wang 외

Human emotional expression is inherently dynamic, complex, and fluid, characterized by smooth transitions in intensity throughout verbal communication. However, the modeling of such intensity fluctuations has been largel…

Talking Head Generation

FlowPortrait: Reinforcement Learning for Audio-Driven Portrait Video Generation

2026-02-25 · Weiting Tan, Andy T. Liu, Ming Tu, Xinghua Qu 외 arxiv

Generating realistic talking-head videos remains challenging due to persistent issues such as imperfect lip synchronization, unnatural motion, and evaluation metrics that correlate poorly with human perception. We propos…

Reinforcement LearningVideo Generation

Semantic Processing of Political Words in Naturalistic Information Differs by Political Orientation

2023-05-12 · Shuhei Kitamura, Aya S. Ihara

Worldviews may differ significantly according to political orientation. Even a single word can have a completely different meaning depending on political orientation. However, no direct evidence has been obtained on diff…

Controllable Talking Face Generation by Implicit Facial Keypoints Editing

2024-06-05 · Dong Zhao, Jiaying Shi, Wenjun Li, Shudong Wang 외

Audio-driven talking face generation has garnered significant interest within the domain of digital human research. Existing methods are encumbered by intricate model architectures that are intricately dependent on each …

Face GenerationTalking Face Generation

GoHD: Gaze-oriented and Highly Disentangled Portrait Animation with Rhythmic Poses and Realistic Expression

2024-12-12 · Ziqi Zhou, Weize Quan, Hailin Shi, Wei Li 외

Audio-driven talking head generation necessitates seamless integration of audio and visual data amidst the challenges posed by diverse input portraits and intricate correlations between audio and facial motions. In respo…

DisentanglementPortrait AnimationTalking Head Generation