paper-with-me

Papers

VIVECaption: A Split Approach to Caption Quality Improvement

2026-03-08 · Varun Ananth, Baqiao Liu, Haoran Cai arxiv

Caption quality has emerged as a critical bottleneck in training high-quality text-to-image (T2I) and text-to-video (T2V) generative models. While visual language models (VLMs) are commonly deployed to generate captions from visual data, they suffer from hallucinations, poor compositional reasoning, and limited fine-grained understanding, resulting in misaligned image-caption pairs that degrade downstream model performance. This technical report introduces VIVECaption, a systematic two-sided approach to caption quality improvement. We first establish a comprehensive taxonomy of caption evaluation metrics, distinguishing between "universal" and "instance-grounded" metrics, with the ultimate goal of showcasing the use-cases and tradeoffs between different caption quality metrics. We then use this language to describe our two-sided approach to caption quality improvement: (1) a gold-standard dataset creation methodology using stratified sampling and (2) a model alignment strategy encompassing context alignment and parameter-level finetuning using SFT. We demonstrate our methodology on open-source models, focusing on structured caption formats that enable better parsing and downstream utilization. We ultimately show that using a finetuned character detection model in an image captioning pipeline significantly improves holistic image-caption alignment quality. Our work addresses the growing need for high-quality "vegan" training data in enterprise AI development, providing practical solutions for teams seeking to improve caption-image alignment without relying on potentially copyright-protected web-scraped content.

📄 PDF Abstract BibTeX arXiv:2603.07401

Code (0)

등록된 구현이 없습니다.

Tasks

Image Captioning

Similar Papers 제목 키워드 기반

Dual-Stream Transformer for Generic Event Boundary Captioning

2022-07-07 · Xin Gu, Hanhua Ye, Guang Chen, YuFei Wang 외

This paper describes our champion solution for the CVPR2022 Generic Event Boundary Captioning (GEBC) competition. GEBC requires the captioning model to have a comprehension of instantaneous status changes around the give…

Boundary CaptioningVideo Captioning

Improving Explicit Spatial Relationships in Text-to-Image Generation through an Automatically Derived Dataset

2024-03-01 · Ander Salaberria, Gorka Azkune, Oier Lopez de Lacalle, Aitor Soroa 외

Existing work has observed that current text-to-image systems do not accurately reflect explicit spatial relations between objects such as 'left of' or 'below'. We hypothesize that this is because explicit spatial relati…

Image CaptioningImage GenerationText to Image GenerationText-to-Image Generation

Koala-36M: A Large-scale Video Dataset Improving Consistency between Fine-grained Conditions and Video Content

2024-10-10 · CVPR 2025 1 · Qiuheng Wang, Yukai Shi, Jiarong Ou, Rui Chen 외

As visual generation technologies continue to advance, the scale of video datasets has expanded rapidly, and the quality of these datasets is critical to the performance of video generation models. We argue that temporal…

Video AlignmentVideo Generation

Show, Interpret and Tell: Entity-aware Contextualised Image Captioning in Wikipedia

2022-09-21 · Khanh Nguyen, Ali Furkan Biten, Andres Mafla, Lluis Gomez 외

Humans exploit prior knowledge to describe images, and are able to adapt their explanation to specific contextual information, even to the extent of inventing plausible explanations when contextual information and images…

ArticlesImage Captioning

Panda-70M: Captioning 70M Videos with Multiple Cross-Modality Teachers

2024-02-29 · CVPR 2024 1 · Tsai-Shien Chen, Aliaksandr Siarohin, Willi Menapace, Ekaterina Deyneka 외

The quality of the data and annotation upper-bounds the quality of a downstream model. While there exist large text corpora and image-text pairs, high-quality video-text data is much harder to collect. First of all, manu…

RetrievalText RetrievalVideo CaptioningVideo Description+1