paper-with-me

홈 › Papers

PodBench: A Comprehensive Benchmark for Instruction-Aware Audio-Oriented Podcast Script Generation

2026-01-21 · Chenning Xu, Mao Zheng, Mingyu Zheng, Mingyang Song arxiv

Podcast script generation requires LLMs to synthesize structured, context-grounded dialogue from diverse inputs, yet systematic evaluation resources for this task remain limited. To bridge this gap, we introduce PodBench, a benchmark comprising 800 samples with inputs up to 21K tokens and complex multi-speaker instructions. We propose a multifaceted evaluation framework that integrates quantitative constraints with LLM-based quality assessment. Extensive experiments reveal that while proprietary models generally excel, open-source models equipped with explicit reasoning demonstrate superior robustness in handling long contexts and multi-speaker coordination compared to standard baselines. However, our analysis uncovers a persistent divergence where high instruction following does not guarantee high content substance. PodBench offers a reproducible testbed to address these challenges in long-form, audio-centric generation.

📄 PDF Abstract BibTeX arXiv:2601.14903

Code (0)

등록된 구현이 없습니다.

Tasks

Instruction Following

Similar Papers 제목 키워드 기반

Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models

2026-07-15 · Chun-Yi Kuan, Siwon Kim, Byeonggeun Kim, Suyoun Kim 외 arxiv

Recent text-to-audio models generate high-quality audio, but often fail to follow instructions involving multiple sound events and temporal order. This gap arises because existing evaluation and training signals mainly e…

Instruction Following

CMI-Bench: A Comprehensive Benchmark for Evaluating Music Instruction Following

2025-06-14 · Yinghao Ma, Siyou Li, Juntao Yu, Emmanouil Benetos 외

Recent advances in audio-text large language models (LLMs) have opened new possibilities for music understanding and generation. However, existing benchmarks are limited in scope, often relying on simplified tasks or mul…

Beat TrackingGenre classificationInformation RetrievalInstruction Following+5

ISA-Bench: Benchmarking Instruction Sensitivity for Large Audio Language Models

2025-10-27 · Bohan Li, Wenbin Huang, Yuhang Qiu, Yiwei Guo 외 arxiv

Large Audio Language Models (LALMs), which couple acoustic perception with large language models (LLMs) to extract and understand diverse information from audio, have attracted intense interest from both academic and ind…

INSPIRE: A Benchmark for Instruction-Aware Speech Retrieval

2026-08-17 · Chen-An Li, Hung-yi Lee arxiv

Existing speech retrieval systems rely on fixed similarity matching and cannot adapt to diverse user intents. We introduce INSPIRE, the first benchmark for instruction-aware speech retrieval, in which natural-language in…

Semantic Retrieval

AVE-Compass: Towards Holistic Evaluation for Audio-Video Editing Abilities

2026-07-17 · Yuqing Wen, Yukai Huang, Qianqian Xie, Jiangtao Wu 외 hf

While instruction-based video editing has advanced rapidly, real-world videos contain tightly coupled audio and visual signals, and editing one modality often requires coordinated changes in the other. Existing benchmark…

Instruction Following