paper-with-me

홈 › Papers

T2AV-Compass: Towards Unified Evaluation for Text-to-Audio-Video Generation

2025-12-24 · Zhe Cao, Tao Wang, Jiaming Wang, Yanghai Wang, Yuanxing Zhang, Jiahao Wang, Jialu Chen, Miao Deng, Yubin Guo, Chenxi Liao, Yize Zhang, Zhaoxiang Zhang, Jiaheng Liu arxiv

Text-to-Audio-Video (T2AV) generation aims to synthesize temporally coherent video and semantically synchronized audio from natural language, yet its evaluation remains fragmented, often relying on unimodal metrics or narrowly scoped benchmarks that fail to capture cross-modal alignment, instruction following, and perceptual realism under complex prompts. To address this limitation, we present T2AV-Compass, a unified benchmark for comprehensive evaluation of T2AV systems, consisting of 500 diverse and complex prompts constructed via a taxonomy-driven pipeline to ensure semantic richness and physical plausibility. Besides, T2AV-Compass introduces a dual-level evaluation framework that integrates objective signal-level metrics for video quality, audio quality, and cross-modal alignment with a subjective MLLM-as-a-Judge protocol for instruction following and realism assessment. Extensive evaluation of 11 representative T2AVsystems reveals that even the strongest models fall substantially short of human-level realism and cross-modal consistency, with persistent failures in audio realism, fine-grained synchronization, instruction following, etc. These results indicate significant improvement room for future models and highlight the value of T2AV-Compass as a challenging and diagnostic testbed for advancing text-to-audio-video generation.

📄 PDF Abstract BibTeX arXiv:2512.21094

Code (0)

등록된 구현이 없습니다.

Tasks

Instruction FollowingVideo Generation

Similar Papers 제목 키워드 기반

LongAV-Compass: Towards Unified Evaluation of Minute-Scale Audio-Visual Generation Across T2AV, I2AV, and V2AV

2026-05-25 · Tengfei Liu, Yang Shi, Xuanyu Zhu, Jiafu Tang 외 arxiv

Audio-visual generation is rapidly advancing from short clips to minute-long content, while existing evaluation protocols remain largely confined to short-form settings. Existing benchmarks primarily focus on 5--10 secon…

MultiRef-Compass: Towards Comprehensive Evaluation of Multi-Reference-to-Audio-Video Generation

2026-07-15 · Xiaohan Zhang, Yuqing Wen, Junlin Chen, Yuqi Tang 외 hf

Multi-reference-to-audio-video (MR2AV) generation aims to generate coherent audio-video content conditioned on multiple references and textual instructions. Existing benchmarks mainly focus on text-driven generation, sin…

Instruction FollowingVideo GenerationVideo Alignment

VidAudio-Bench: Benchmarking V2A and VT2A Generation across Four Audio Categories

2026-04-12 · Qian Zhang, Yuqin Cao, Yixuan Gao, Xiongkuo Min arxiv

Video-to-Audio (V2A) generation is essential for immersive multimedia experiences, yet its evaluation remains underexplored. Existing benchmarks typically assess diverse audio types under a unified protocol, overlooking …

Instruction FollowingAudio Generation

VABench: A Comprehensive Benchmark for Audio-Video Generation

2025-12-10 · Daili Hua, Xizhi Wang, Bohan Zeng, Xinyi Huang 외 arxiv

Recent advances in video generation have been remarkable, enabling models to produce visually compelling videos with synchronized audio. While existing video generation benchmarks provide comprehensive metrics for visual…

Video Generation

From Vision to Audio and Beyond: A Unified Model for Audio-Visual Representation and Generation

2024-09-27 · Kun Su, Xiulong Liu, Eli Shlizerman

Video encompasses both visual and auditory data, creating a perceptually rich experience where these two modalities complement each other. As such, videos are a valuable type of media for the investigation of the interpl…

Audio ClassificationAudio GenerationRepresentation Learning