paper-with-me

홈 › Papers

AVoCaDO: An Audiovisual Video Captioner Driven by Temporal Orchestration

2025-10-12 · Xinlong Chen, Yue Ding, Weihong Lin, Jingyun Hua, Linli Yao, Yang Shi, Bozhou Li, Yuanxing Zhang, Qiang Liu, Pengfei Wan, Liang Wang, Tieniu Tan arxiv

Audiovisual video captioning aims to generate semantically rich descriptions with temporal alignment between visual and auditory events, thereby benefiting both video understanding and generation. In this paper, we present AVoCaDO, a powerful audiovisual video captioner driven by the temporal orchestration between audio and visual modalities. We propose a two-stage post-training pipeline: (1) AVoCaDO SFT, which fine-tunes the model on a newly curated dataset of 107K high-quality, temporally-aligned audiovisual captions; and (2) AVoCaDO GRPO, which leverages tailored reward functions to further enhance temporal coherence and dialogue accuracy while regularizing caption length and reducing collapse. Experimental results demonstrate that AVoCaDO significantly outperforms existing open-source models across four audiovisual video captioning benchmarks, and also achieves competitive performance on the VDC and DREAM-1K benchmark under visual-only settings.

📄 PDF Abstract BibTeX arXiv:2510.10395

Code (0)

등록된 구현이 없습니다.

Tasks

Video Captioning

Similar Papers 제목 키워드 기반

Temporal and Cross-Modal Alignment for Enhanced Audiovisual Video Captioning

2026-07-02 · Chen Zhao, Jiajun Ma, Qilong Huang, Tiehan Fan 외 arxiv

While Multimodal Large Language Models (MLLMs) have advanced video understanding, achieving precise temporal and cross-modal alignment in audiovisual video captioning remains a formidable challenge. Most existing approac…

Relational ReasoningVideo Captioning

Towards Universal Video MLLMs with Attribute-Structured and Quality-Verified Instructions

2026-02-13 · Yunheng Li, Hengrui Zhang, Meng-Hao Guo, Wenzhao Gao 외 arxiv

Universal video understanding requires modeling fine-grained visual and audio information over time in diverse real-world scenarios. However, the performance of existing models is primarily constrained by video-instructi…

Instruction Following

Video ChatCaptioner: Towards Enriched Spatiotemporal Descriptions

2023-04-09 · Jun Chen, Deyao Zhu, Kilichbek Haydarov, Xiang Li 외

Video captioning aims to convey dynamic scenes from videos using natural language, facilitating the understanding of spatiotemporal information within our environment. Although there have been recent advances, generating…

Video Captioning

An Eye for an Ear: Zero-shot Audio Description Leveraging an Image Captioner using Audiovisual Distribution Alignment

2024-10-08 · Hugo Malard, Michel Olvera, Stéphane Lathuiliere, Slim Essid

Multimodal large language models have fueled progress in image captioning. These models, fine-tuned on vast image datasets, exhibit a deep understanding of semantic concepts. In this work, we show that this ability can b…

Audio captioningContrastive LearningImage CaptioningRepresentation Learning+1

TMD-Bench: A Multi-Level Evaluation Paradigm for Music-Dance Co-Generation

2026-05-03 · Xiaoda Yang, Majun Zhang, Changhao Pan, Nick Huang 외 arxiv

Unified audio-visual generation is rapidly gaining industrial and creative relevance, enabling applications in virtual production and interactive media. However, when moving from general audio-video synthesis to music-da…