paper-with-me

홈 › Papers

Comparative Analysis of Audio Feature Extraction for Real-Time Talking Portrait Synthesis

2024-11-20 · Pegah Salehi, Sajad Amouei Sheshkal, Vajira Thambawita, Sushant Gautam, Saeed S. Sabet, Dag Johansen, Michael A. Riegler, Pål Halvorsen

This paper examines the integration of real-time talking-head generation for interviewer training, focusing on overcoming challenges in Audio Feature Extraction (AFE), which often introduces latency and limits responsiveness in real-time applications. To address these issues, we propose and implement a fully integrated system that replaces conventional AFE models with Open AI's Whisper, leveraging its encoder to optimize processing and improve overall system efficiency. Our evaluation of two open-source real-time models across three different datasets shows that Whisper not only accelerates processing but also improves specific aspects of rendering quality, resulting in more realistic and responsive talking-head interactions. These advancements make the system a more effective tool for immersive, interactive training applications, expanding the potential of AI-driven avatars in interviewer training.

📄 PDF Abstract BibTeX arXiv:2411.13209

Code (1)

pegahs1993/whisper-afe-talkingheadsgen 공식 구현 pytorch

Tasks

Talking Head Generation

Similar Papers 제목 키워드 기반

Fractal Dimension Pattern Based Multiresolution Analysis for Rough Estimator of Person-Dependent Audio Emotion Recognition

2016-07-01 · Miao Cheng, Ah Chung Tsoi

As a general means of expression, audio analysis and recognition has attracted much attentions for its wide applications in real-life world. Audio emotion recognition (AER) attempts to understand emotional states of huma…

Audio Emotion RecognitionEmotion Recognition

Comparative Analysis of Mel-Frequency Cepstral Coefficients and Wavelet Based Audio Signal Processing for Emotion Detection and Mental Health Assessment in Spoken Speech

2024-12-12 · Idoko Agbo, Dr Hoda El-Sayed, M. D Kamruzzan Sarker

The intersection of technology and mental health has spurred innovative approaches to assessing emotional well-being, particularly through computational techniques applied to audio data analysis. This study explores the …

Audio Signal ProcessingData Augmentation

Fundamental Survey on Neuromorphic Based Audio Classification

2025-02-20 · Amlan Basu, Pranav Chaudhari, Gaetano Di Caterina

Audio classification is paramount in a variety of applications including surveillance, healthcare monitoring, and environmental analysis. Traditional methods frequently depend on intricate signal processing algorithms an…

Audio ClassificationClassificationComputational EfficiencySurvey

Music Genre Classification Using Machine Learning Techniques

2025-09-01 · Alokit Mishra, Ryyan Akhtar arxiv

This paper presents a comparative analysis of machine learning methodologies for automatic music genre classification. We evaluate the performance of classical classifiers, including Support Vector Machines (SVM) and ens…

Genre classificationFeature Engineering

Bridging Biological Hearing and Neuromorphic Computing: End-to-End Time-Domain Audio Signal Processing with Reservoir Computing

2026-03-25 · Rinku Sebastian, Simon O'Keefe, Martin Trefzer arxiv

Despite the advancements in cutting-edge technologies, audio signal processing continues to pose challenges and lacks the precision of a human speech processing system. To address these challenges, we propose a novel app…

Speech Recognition