paper-with-me

홈 › Papers

CAPEval: A Decoupled Caption Evaluation across Understanding and Generation

2026-08-03 · Zhipeng Liu, Haochen Wang, Zhaoxiang Zhang hf

Captions serve as a primary supervision signal for both multimodal understanding and text-to-image generation. However, previous evaluations treat the caption quality as a single scalar objective, which conflates two distinct properties: (1) how much visual information a caption covers and (2) how reliably the image supports its stated claims. To this end, we design a decoupled caption evaluation benchmark, CAPEval (Coverage And Precision Evaluation), with human-written ground-truth captions and human-verified atomic checklist items. Specifically, CAPEval decomposes caption quality into Coverage and Precision. The former quantifies how thoroughly a caption covers ground-truth factual content, while the latter reflects the factual correctness rate of all claims expressed in the caption. We select 10 captioners and further conduct controlled downstream end-to-end experiments with them from four model families, where the caption source is the only variable. Empirically, we find a consistent task-dependent dissociation: Coverage serves as the stronger correlate for understanding performance, whereas Precision acts as the dominant predictor for generation performance. This decoupled evaluation paradigm not only delivers a more fine-grained diagnosis of caption quality, but also offers actionable guidance for selecting and optimizing captioners tailored to different downstream tasks.

📄 PDF Abstract BibTeX arXiv:2608.02589

Code (0)

등록된 구현이 없습니다.

Tasks

Text-to-Image Generation

Similar Papers 제목 키워드 기반

AnyCap Project: A Unified Framework, Dataset, and Benchmark for Controllable Omni-modal Captioning

2025-07-17 · Yiming Ren, Zhiqiang Lin, Yu Li, Gao Meng 외

Controllable captioning is essential for precise multimodal alignment and instruction following, yet existing models often lack fine-grained control and reliable evaluation protocols. To address this gap, we present the …

Instruction Following

Fine-grained Image Captioning with CLIP Reward

2022-05-26 · Findings (NAACL) 2022 7 · Jaemin Cho, Seunghyun Yoon, Ajinkya Kale, Franck Dernoncourt 외

Modern image captioning models are usually trained with text similarity objectives. However, since reference captions in public datasets often describe the most salient common objects, models trained with text similarity…

Caption GenerationDescriptiveImage CaptioningImage Retrieval+3

Progress-Aware Video Frame Captioning

2024-12-03 · CVPR 2025 1 · Zihui Xue, Joungbin An, Xitong Yang, Kristen Grauman

While image captioning provides isolated descriptions for individual images, and video captioning offers one single narrative for an entire video clip, our work explores an important middle ground: progress-aware video c…

Image CaptioningVideo CaptioningVideo Understanding

Empowering Long-form Omni-modal Understanding with Robust Audio Perception

2026-07-11 · Kaiying Yan, Luoyi Sun, Xiao Zhou, Weidi Xie arxiv

Recent advances in large-scale multimodal models have drivenremarkable progress in vision-language tasks; however, comprehensiveomni-modal understanding remains under-explored, largely due to thescarcity of datasets with…

Video Captioning

CapRL++: Unified Reinforcement Learning with Verifiable Rewards for Dense Image and Video Captioning

2026-06-08 · Penghui Yang, Long Xing, Xiaoyi Dong, Yuhang Zang 외 arxiv

Image and video captioning are fundamental tasks that bridge the visual and linguistic domains, playing a critical role in pre-training Large Vision-Language Models (LVLMs). Current state-of-the-art captioning models are…

Reinforcement LearningVideo CaptioningDense Captioning