paper-with-me

홈 › Papers

Audio Flamingo Next: Next-Generation Open Audio-Language Models for Speech, Sound, and Music

2026-04-13 · Sreyan Ghosh, Arushi Goel, Kaousheik Jayakumar, Lasha Koroshinadze, Nishit Anand, Zhifeng Kong, Siddharth Gururani, Sang-gil Lee, Jaehyeon Kim, Aya Aljafari, Chao-Han Huck Yang, Sungwon Kim, Ramani Duraiswami, Dinesh Manocha, Mohammad Shoeybi, Bryan Catanzaro, Ming-Yu Liu, Wei Ping arxiv

We present Audio Flamingo Next (AF-Next), the next-generation and most capable large audio-language model in the Audio Flamingo series, designed to advance understanding and reasoning over speech, environmental sounds and music. Compared to Audio Flamingo 3, AF-Next introduces: (i) a stronger foundational audio-language model that significantly improves accuracy across diverse audio understanding tasks; (ii) scalable strategies for constructing large-scale audio understanding and reasoning data beyond existing academic benchmarks; (iii) support for long and complex audio inputs up to 30 minutes; and (iv) Temporal Audio Chain-of-Thought, a new reasoning paradigm that explicitly grounds intermediate reasoning steps to timestamps in long audio, enabling fine-grained temporal alignment and improved interpretability. To enable these capabilities, we first conduct a systematic analysis of Audio Flamingo 3 to identify key gaps in audio understanding and reasoning. We then curate and scale new large-scale datasets totaling over 1 million hours to address these limitations and expand the existing AudioSkills-XL, LongAudio-XL, AF-Think and AF-Chat datasets. AF-Next is trained using a curriculum-based strategy spanning pre-training, mid-training and post-training stages. Extensive experiments across 20 audio understanding and reasoning benchmarks, including challenging long-audio tasks, show that AF-Next outperforms similarly sized open models by large margins and remains highly competitive with and sometimes surpasses, much larger open-weight and closed models. Beyond benchmark performance, AF-Next exhibits strong real-world utility and transfers well to unseen tasks, highlighting its robustness and generalization ability. In addition to all data, code and methods, we open-source 3 variants of AF-Next, including AF-Next-Instruct, AF-Next-Think and AF-Next-Captioner.

📄 PDF Abstract BibTeX arXiv:2604.10905

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Music Flamingo: Scaling Music Understanding in Audio Language Models

2025-11-13 · Sreyan Ghosh, Arushi Goel, Lasha Koroshinadze, Sang-gil Lee 외 arxiv

We introduce Music Flamingo, a novel large audio-language model designed to advance music (including song) understanding in foundational audio models. While audio-language research has progressed rapidly, music remains c…

Reinforcement Learning

Audio Flamingo 2: An Audio-Language Model with Long-Audio Understanding and Expert Reasoning Abilities

2025-03-06 · Sreyan Ghosh, Zhifeng Kong, Sonal Kumar, S Sakshi 외

Understanding and reasoning over non-speech sounds and music are crucial for both humans and AI agents to interact effectively with their environments. In this paper, we introduce Audio Flamingo 2 (AF2), an Audio-Languag…

Audio captioningLanguage ModelingLanguage ModellingQuestion Answering+1

Audio-Visual Flamingo: Open Audio-Visual Intelligence for Long and Complex Videos

2026-07-17 · Sreyan Ghosh, Arushi Goel, Kaousheik Jayakumar, Lasha Koroshinadze 외 arxiv

We present Audio-Visual Flamingo (AV-Flamingo), a fully open state-of-the-art audio-visual large language model (AV-LLM) for joint understanding and reasoning over audio, images, and long-form videos. Unlike prior AV-LLM…

Visual Reasoning

Audio Flamingo: A Novel Audio Language Model with Few-Shot Learning and Dialogue Abilities

2024-02-02 · Zhifeng Kong, Arushi Goel, Rohan Badlani, Wei Ping 외

Augmenting large language models (LLMs) to understand audio -- including non-speech sounds and non-verbal speech -- is critically important for diverse real-world applications of LLMs. In this paper, we propose Audio Fla…

Acoustic Scene ClassificationAudio captioningFew-Shot LearningIn-Context Learning+5

Audio Flamingo 3: Advancing Audio Intelligence with Fully Open Large Audio Language Models

2025-07-10 · Arushi Goel, Sreyan Ghosh, Jaehyeon Kim, Sonal Kumar 외

We present Audio Flamingo 3 (AF3), a fully open state-of-the-art (SOTA) large audio-language model that advances reasoning and understanding across speech, sound, and music. AF3 introduces: (i) AF-Whisper, a unified audi…

Language ModelingLanguage ModellingRepresentation Learning