paper-with-me

홈 › Papers

EVA-GAN: Enhanced Various Audio Generation via Scalable Generative Adversarial Networks

2024-01-31 · Shijia Liao, Shiyi Lan, Arun George Zachariah

The advent of Large Models marks a new era in machine learning, significantly outperforming smaller models by leveraging vast datasets to capture and synthesize complex patterns. Despite these advancements, the exploration into scaling, especially in the audio generation domain, remains limited, with previous efforts didn't extend into the high-fidelity (HiFi) 44.1kHz domain and suffering from both spectral discontinuities and blurriness in the high-frequency domain, alongside a lack of robustness against out-of-domain data. These limitations restrict the applicability of models to diverse use cases, including music and singing generation. Our work introduces Enhanced Various Audio Generation via Scalable Generative Adversarial Networks (EVA-GAN), yields significant improvements over previous state-of-the-art in spectral and high-frequency reconstruction and robustness in out-of-domain data performance, enabling the generation of HiFi audios by employing an extensive dataset of 36,000 hours of 44.1kHz audio, a context-aware module, a Human-In-The-Loop artifact measurement toolkit, and expands the model to approximately 200 million parameters. Demonstrations of our work are available at https://double-blind-eva-gan.cc.

📄 PDF Abstract BibTeX arXiv:2402.00892

Code (1)

fishaudio/vocoder pytorch

Tasks

Audio GenerationSpeech Synthesis

Similar Papers 제목 키워드 기반

Model-Guided Dual-Role Alignment for High-Fidelity Open-Domain Video-to-Audio Generation

2025-10-28 · Kang Zhang, Trung X. Pham, Suyeon Lee, Axi Niu 외 arxiv

We present MGAudio, a novel flow-based framework for open-domain video-to-audio generation, which introduces model-guided dual-role alignment as a central design principle. Unlike prior approaches that rely on classifier…

Audio Generation

Make-An-Audio: Text-To-Audio Generation with Prompt-Enhanced Diffusion Models

2023-01-30 · Rongjie Huang, Jiawei Huang, Dongchao Yang, Yi Ren 외

Large-scale multimodal generative modeling has created milestones in text-to-image and text-to-video generation. Its application to audio still lags behind for two main reasons: the lack of large-scale datasets with high…

Audio GenerationText-to-Video GenerationVideo Generation

Enhanced Generative Machine Listener

2025-09-25 · Vishnu Raj, Gouthaman KV, Shiv Gehlot, Lars Villemoes 외 arxiv

We present GMLv2, a reference-based model designed for the prediction of subjective audio quality as measured by MUSHRA scores. GMLv2 introduces a Beta distribution-based loss to model the listener ratings and incorporat…

AI-Based Sound Effect Generation: A Narrative Review of Generative Models Across Input Modalities

2026-08-04 · Sandy Abdo, Bill Kapralos, Priyamvada Tripathi, KC Collins 외 arxiv

Sound effects play a crucial role in conveying actions, events, and environmental cues across digital applications, often requiring a high degree of variation and contextual adaptability. Artificial intelligence (AI)-dri…

MultiRef-Compass: Towards Comprehensive Evaluation of Multi-Reference-to-Audio-Video Generation

2026-07-15 · Xiaohan Zhang, Yuqing Wen, Junlin Chen, Yuqi Tang 외 hf

Multi-reference-to-audio-video (MR2AV) generation aims to generate coherent audio-video content conditioned on multiple references and textual instructions. Existing benchmarks mainly focus on text-driven generation, sin…

Instruction FollowingVideo GenerationVideo Alignment