paper-with-me

Papers

Multi Modal Adaptive Normalization for Audio to Video Generation

2020-12-14 · Neeraj Kumar, Srishti Goel, Ankur Narang, Brejesh lall

Speech-driven facial video generation has been a complex problem due to its multi-modal aspects namely audio and video domain. The audio comprises lots of underlying features such as expression, pitch, loudness, prosody(speaking style) and facial video has lots of variability in terms of head movement, eye blinks, lip synchronization and movements of various facial action units along with temporal smoothness. Synthesizing highly expressive facial videos from the audio input and static image is still a challenging task for generative adversarial networks. In this paper, we propose a multi-modal adaptive normalization(MAN) based architecture to synthesize a talking person video of arbitrary length using as input: an audio signal and a single image of a person. The architecture uses the multi-modal adaptive normalization, keypoint heatmap predictor, optical flow predictor and class activation map[58] based layers to learn movements of expressive facial components and hence generates a highly expressive talking-head video of the given person. The multi-modal adaptive normalization uses the various features of audio and video such as Mel spectrogram, pitch, energy from audio signals and predicted keypoint heatmap/optical flow and a single image to learn the respective affine parameters to generate highly expressive video. Experimental evaluation demonstrates superior performance of the proposed method as compared to Realistic Speech-Driven Facial Animation with GANs(RSDGAN) [53], Speech2Vid [10], and other approaches, on multiple quantitative metrics including: SSIM (structural similarity index), PSNR (peak signal to noise ratio), CPBD (image sharpness), WER(word error rate), blinks/sec and LMD(landmark distance). Further, qualitative evaluation and Online Turing tests demonstrate the efficacy of our approach.

📄 PDF Abstract BibTeX arXiv:2012.07304

Code (0)

등록된 구현이 없습니다.

Tasks

Optical Flow EstimationSSIMVideo Generation

Methods 이 논문이 사용한 방법론

Heatmap 설명 없음

Similar Papers 제목 키워드 기반

A New View of Multi-modal Language Analysis: Audio and Video Features as Text ``Styles''

2021-04-01 · EACL 2021 2 · Zhongkai Sun, Prathusha K Sarma, YIngyu Liang, William Sethares

Imposing the style of one image onto another is called style transfer. For example, the style of a Van Gogh painting might be imposed on a photograph to yield an interesting hybrid. This paper applies the adaptive normal…

Emotion RecognitionSentiment AnalysisStyle Transfer

One Shot Audio to Animated Video Generation

2021-02-19 · Neeraj Kumar, Srishti Goel, Ankur Narang, Brejesh lall 외

We consider the challenging problem of audio to animated video generation. We propose a novel method OneShotAu2AV to generate an animated video of arbitrary length using an audio clip and a single unseen image of a perso…

Video Generation

Multi-Modulation Network for Audio-Visual Event Localization

2021-08-26 · Hao Wang, Zheng-Jun Zha, Liang Li, Xuejin Chen 외

We study the problem of localizing audio-visual events that are both audible and visible in a video. Existing works focus on encoding and aligning audio and visual features at the segment level while neglecting informati…

audio-visual event localization

AccKV: Towards Efficient Audio-Video LLMs Inference via Adaptive-Focusing and Cross-Calibration KV Cache Optimization

2025-11-14 · Zhonghua Jiang, Kui Chen, Kunxi Li, Keting Yin 외 arxiv

Recent advancements in Audio-Video Large Language Models (AV-LLMs) have enhanced their capabilities in tasks like audio-visual question answering and multimodal dialog systems. Video and audio introduce an extended tempo…

Audio-visual Question AnsweringComputational Efficiency

An Audio-centric Multi-task Learning Framework for Streaming Ads Targeting on Spotify

2025-06-23 · Shivam Verma, Vivian Chen, Darren Mei

Spotify, a large-scale multimedia platform, attracts over 675 million monthly active users who collectively consume millions of hours of music, podcasts, audiobooks, and video content. This diverse content consumption pa…

Click-Through Rate PredictionMixture-of-ExpertsMulti-Task Learning