paper-with-me

홈 › Papers

Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models

2026-07-15 · Chun-Yi Kuan, Siwon Kim, Byeonggeun Kim, Suyoun Kim, Bo-Ru Lu, Qingming Tang, Ankur Gandhe, Hung-yi Lee, Chieh-Chi Kao, Chao Wang arxiv

Recent text-to-audio models generate high-quality audio, but often fail to follow instructions involving multiple sound events and temporal order. This gap arises because existing evaluation and training signals mainly emphasize global similarity or perceptual quality, with limited supervision on instruction-level correctness. We propose an instruction-level framework that uses audio-aware large language models (ALLMs) as fine-grained judges to verify target event presence and temporal relations in generated audio. After validating ALLM judgments on benchmarks and through human verification, we use their feedback to construct preference pairs for direct preference optimization. We further introduce S3Bench, a narrative benchmark for evaluating multi-event temporal instruction following. Experiments show that our method improves event completeness, temporal ordering, and joint instruction-following accuracy across existing benchmarks and S3Bench, while maintaining audio quality.

📄 PDF Abstract BibTeX arXiv:2607.13408

Code (0)

등록된 구현이 없습니다.

Tasks

Instruction Following

Similar Papers 제목 키워드 기반

Towards Universal Video MLLMs with Attribute-Structured and Quality-Verified Instructions

2026-02-13 · Yunheng Li, Hengrui Zhang, Meng-Hao Guo, Wenzhao Gao 외 arxiv

Universal video understanding requires modeling fine-grained visual and audio information over time in diverse real-world scenarios. However, the performance of existing models is primarily constrained by video-instructi…

Instruction Following

T2AV-Compass: Towards Unified Evaluation for Text-to-Audio-Video Generation

2025-12-24 · Zhe Cao, Tao Wang, Jiaming Wang, Yanghai Wang 외 arxiv

Text-to-Audio-Video (T2AV) generation aims to synthesize temporally coherent video and semantically synchronized audio from natural language, yet its evaluation remains fragmented, often relying on unimodal metrics or na…

Instruction FollowingVideo Generation

FusionAudio-1.2M: Towards Fine-grained Audio Captioning with Multimodal Contextual Fusion

2025-06-01 · Shunian Chen, Xinyuan Xie, Zheshu Chen, Liyan Zhao 외

High-quality, large-scale audio captioning is crucial for advancing audio understanding, yet current automated methods often generate captions that lack fine-grained detail and contextual accuracy, primarily due to their…

Audio captioningCaption GenerationInstruction FollowingLarge Language Model

InstructAny2Pix: Flexible Visual Editing via Multimodal Instruction Following

2023-12-11 · Shufan Li, Harkanwar Singh, Aditya Grover

The ability to provide fine-grained control for generating and editing visual imagery has profound implications for computer vision and its applications. Previous works have explored extending controllability in two dire…

DecoderInstruction Following

IFEval-Audio: Benchmarking Instruction-Following Capability in Audio-based Large Language Models

2025-05-22 · Yiming Gao, Bin Wang, Chengwei Wei, Shuo Sun 외

Large language models (LLMs) have demonstrated strong instruction-following capabilities in text-based tasks. However, this ability often deteriorates in multimodal models after alignment with non-text modalities such as…

BenchmarkingInstruction Following