paper-with-me

Papers

AudioGPT: Understanding and Generating Speech, Music, Sound, and Talking Head

2023-04-25 · Rongjie Huang, Mingze Li, Dongchao Yang, Jiatong Shi, Xuankai Chang, Zhenhui Ye, Yuning Wu, Zhiqing Hong, Jiawei Huang, Jinglin Liu, Yi Ren, Zhou Zhao, Shinji Watanabe

Large language models (LLMs) have exhibited remarkable capabilities across a variety of domains and tasks, challenging our understanding of learning and cognition. Despite the recent success, current LLMs are not capable of processing complex audio information or conducting spoken conversations (like Siri or Alexa). In this work, we propose a multi-modal AI system named AudioGPT, which complements LLMs (i.e., ChatGPT) with 1) foundation models to process complex audio information and solve numerous understanding and generation tasks; and 2) the input/output interface (ASR, TTS) to support spoken dialogue. With an increasing demand to evaluate multi-modal LLMs of human intention understanding and cooperation with foundation models, we outline the principles and processes and test AudioGPT in terms of consistency, capability, and robustness. Experimental results demonstrate the capabilities of AudioGPT in solving AI tasks with speech, music, sound, and talking head understanding and generation in multi-round dialogues, which empower humans to create rich and diverse audio content with unprecedented ease. Our system is publicly available at \url{https://github.com/AIGC-Audio/AudioGPT}.

📄 PDF Abstract BibTeX arXiv:2304.12995

Code (1)

aigc-audio/audiogpt 공식 구현 pytorch

Methods 이 논문이 사용한 방법론

Test 설명 없음

Similar Papers 제목 키워드 기반

Source Separation & Automatic Transcription for Music

2024-12-09 · Bradford Derby, Lucas Dunker, Samarth Galchar, Shashank Jarmale 외

Source separation is the process of isolating individual sounds in an auditory mixture of multiple sounds [1], and has a variety of applications ranging from speech enhancement and lyric transcription [2] to digital audi…

Music TranscriptionSpeech Enhancement

Audio Flamingo 3: Advancing Audio Intelligence with Fully Open Large Audio Language Models

2025-07-10 · Arushi Goel, Sreyan Ghosh, Jaehyeon Kim, Sonal Kumar 외

We present Audio Flamingo 3 (AF3), a fully open state-of-the-art (SOTA) large audio-language model that advances reasoning and understanding across speech, sound, and music. AF3 introduces: (i) AF-Whisper, a unified audi…

Language ModelingLanguage ModellingRepresentation Learning

SEE-2-SOUND: Zero-Shot Spatial Environment-to-Spatial Sound

2024-06-06 · Rishit Dagli, Shivesh Prakash, Robert Wu, Houman Khosravani

Generating combined visual and auditory sensory experiences is critical for the consumption of immersive content. Recent advances in neural generative models have enabled the creation of high-resolution content across mu…

Audio Generation

AudioCapBench: Quick Evaluation on Audio Captioning across Sound, Music, and Speech

2026-02-27 · Jielin Qiu, Jianguo Zhang, Zixiang Chen, Liangwei Yang 외 arxiv

We introduce AudioCapBench, a benchmark for evaluating audio captioning capabilities of large multimodal models. \method covers three distinct audio domains, including environmental sound, music, and speech, with 1,000 c…

Audio captioning

Transferring neural speech waveform synthesizers to musical instrument sounds generation

2019-10-27 · Yi Zhao, Xin Wang, Lauri Juvela, Junichi Yamagishi

Recent neural waveform synthesizers such as WaveNet, WaveGlow, and the neural-source-filter (NSF) model have shown good performance in speech synthesis despite their different methods of waveform generation. The similari…

Audio GenerationAudio SynthesisSpeech SynthesisZero-Shot Learning