paper-with-me

홈 › Papers

Expertized Caption Auto-Enhancement for Video-Text Retrieval

2025-02-05 · Baoyao Yang, Junxiang Chen, Wanyun Li, Wenbin Yao, Yang Zhou

Video-text retrieval has been stuck in the information mismatch caused by personalized and inadequate textual descriptions of videos. The substantial information gap between the two modalities hinders an effective cross-modal representation alignment, resulting in ambiguous retrieval results. Although text rewriting methods have been proposed to broaden text expressions, the modality gap remains significant, as the text representation space is hardly expanded with insufficient semantic enrichment.Instead, this paper turns to enhancing visual presentation, bridging video expression closer to textual representation via caption generation and thereby facilitating video-text matching.While multimodal large language models (mLLM) have shown a powerful capability to convert video content into text, carefully crafted prompts are essential to ensure the reasonableness and completeness of the generated captions. Therefore, this paper proposes an automatic caption enhancement method that improves expression quality and mitigates empiricism in augmented captions through self-learning.Additionally, an expertized caption selection mechanism is designed and introduced to customize augmented captions for each video, further exploring the utilization potential of caption augmentation.Our method is entirely data-driven, which not only dispenses with heavy data collection and computation workload but also improves self-adaptability by circumventing lexicon dependence and introducing personalized matching. The superiority of our method is validated by state-of-the-art results on various benchmarks, specifically achieving Top-1 recall accuracy of 68.5% on MSR-VTT, 68.1% on MSVD, and 62.0% on DiDeMo. Our code is publicly available at https://github.com/CaryXiang/ECA4VTR.

📄 PDF Abstract BibTeX arXiv:2502.02885

Code (1)

caryxiang/eca4vtr 공식 구현 pytorch

Tasks

Caption GenerationRetrievalText RetrievalVideo-Text Retrieval

Similar Papers 제목 키워드 기반

CAPE-T2V: Captioner-Anchored Prompt Enhancement toward Two-Sided Conditioning Alignment in Text-to-Video Generation

2026-08-04 · Yizhuo Jia, Jingyun Hua, Yuanxing Zhang arxiv

Text-to-video (T2V) diffusion transformers (DiTs) are trained with detailed video captions, whereas inference often relies on user prompts rewritten by a prompt enhancer (PE). Prior work has improved generation by optimi…

Text-to-Video Generation

InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption

2024-12-12 · CVPR 2025 1 · Tiehan Fan, Kepan Nan, Rui Xie, Penghao Zhou 외

Text-to-video generation has evolved rapidly in recent years, delivering remarkable results. Training typically relies on video-caption paired data, which plays a crucial role in enhancing generation performance. However…

Text-to-Video GenerationVideo Generation

Explicit Temporal-Semantic Modeling for Dense Video Captioning via Context-Aware Cross-Modal Interaction

2025-11-13 · Mingda Jia, Weiliang Meng, Zenghuang Fu, Yiheng Li 외 arxiv

Dense video captioning jointly localizes and captions salient events in untrimmed videos. Recent methods primarily focus on leveraging additional prior knowledge and advanced multi-task architectures to achieve competiti…

Dense Video CaptioningCross-Modal Retrieval

OVC-Net: Object-Oriented Video Captioning with Temporal Graph and Detail Enhancement

2020-03-08 · Fangyi Zhu, Jenq-Neng Hwang, Zhanyu Ma, Guang Chen 외

Traditional video captioning requests a holistic description of the video, yet the detailed descriptions of the specific objects may not be available. Without associating the moving trajectories, these image-based data-d…

ObjectSentenceVideo Captioning

Narrating the Video: Boosting Text-Video Retrieval via Comprehensive Utilization of Frame-Level Captions

2025-03-07 · CVPR 2025 1 · Chan hur, Jeong-hun Hong, Dong-hun Lee, Dabin Kang 외

In recent text-video retrieval, the use of additional captions from vision-language models has shown promising effects on the performance. However, existing models using additional captions often have struggled to captur…

RetrievalVideo RetrievalVideo Similarity