paper-with-me

홈 › Papers

UALM: Unified Audio Language Model for Understanding, Generation and Reasoning

2025-10-13 · Jinchuan Tian, Sang-gil Lee, Zhifeng Kong, Sreyan Ghosh, Arushi Goel, Chao-Han Huck Yang, Wenliang Dai, Zihan Liu, Hanrong Ye, Shinji Watanabe, Mohammad Shoeybi, Bryan Catanzaro, Rafael Valle, Wei Ping arxiv

Recent advances in the audio language modeling (ALM) domain tackle audio understanding and text-to-audio generation as separate tasks. Very few studies attempt to unify these tasks -- an essential step toward advanced multimodal reasoning. This paper introduces U}nified Audio Language Model (UALM), which aims to unify audio understanding, text-to-audio generation, and multimodal reasoning in a single model. To achieve this goal, we first present UALM-Gen, a text-to-audio language model that directly predicts audio tokens and is comparable to state-of-the-art diffusion-based models. We then demonstrate, using proper data blending, training recipes, and inference techniques, that our single UALM model matches the quality of state-of-the-art specialized models in audio understanding, text-to-audio generation, and text reasoning. Furthermore, we present UALM-Reason, a multimodal reasoning model that utilizes both text and audio in the intermediate thinking steps to facilitate complex generation tasks. To our knowledge, this is the first demonstration in audio research of cross-modal generative reasoning, with its effectiveness confirmed by subjective evaluations.

📄 PDF Abstract BibTeX arXiv:2510.12000

Code (0)

등록된 구현이 없습니다.

Tasks

Multimodal ReasoningAudio Generation

Similar Papers 제목 키워드 기반

VisualMRC: Machine Reading Comprehension on Document Images

2021-01-27 · Ryota Tanaka, Kyosuke Nishida, Sen Yoshida

Recent studies on machine reading comprehension have focused on text-level understanding but have not yet reached the level of human understanding of the visual layout and content of real-world documents. In this study, …

Machine Reading ComprehensionNatural Language UnderstandingQuestion AnsweringReading Comprehension+2

Ming-UniAudio: Speech LLM for Joint Understanding, Generation and Editing with Unified Representation

2025-10-26 · Canxiang Yan, Chunxiang Jin, Dawei Huang, Haibing Yu 외 arxiv

Existing speech models suffer from competing requirements on token representations by understanding and generation tasks. This discrepancy in representation prevents speech language models from performing instruction-bas…

Audio-FLAN: A Preliminary Release

2025-02-23 · Liumeng Xue, Ziya Zhou, Jiahao Pan, Zixuan Li 외

Recent advancements in audio tokenization have significantly enhanced the integration of audio capabilities into large language models (LLMs). However, audio understanding and generation are often treated as distinct tas…

Zero-Shot Learning

Audio-Omni: Extending Multi-modal Understanding to Versatile Audio Generation and Editing

2026-04-12 · Zeyue Tian, Binxin Yang, Zhaoyang Liu, Jiexuan Zhang 외 arxiv

Recent progress in multimodal models has spurred rapid advances in audio understanding, generation, and editing. However, these capabilities are typically addressed by specialized models, leaving the development of a tru…

Audio Generation

Unified-IO 2: Scaling Autoregressive Multimodal Models with Vision, Language, Audio, and Action

2023-12-28 · Jiasen Lu, Christopher Clark, Sangho Lee, Zichen Zhang 외

We present Unified-IO 2, the first autoregressive multimodal model that is capable of understanding and generating image, text, audio, and action. To unify different modalities, we tokenize inputs and outputs -- images, …

DecoderImage GenerationNatural Language Understanding