paper-with-me

Papers

A CLIP-Enhanced Method for Video-Language Understanding

2021-10-14 · Guohao Li, Feng He, Zhifan Feng

This technical report summarizes our method for the Video-And-Language Understanding Evaluation (VALUE) challenge (https://value-benchmark.github.io/challenge\_2021.html). We propose a CLIP-Enhanced method to incorporate the image-text pretrained knowledge into downstream video-text tasks. Combined with several other improved designs, our method outperforms the state-of-the-art by $2.4\%$ ($57.58$ to $60.00$) Meta-Ave score on VALUE benchmark.

📄 PDF Abstract BibTeX arXiv:2110.07137

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Enhanced Motion-Text Alignment for Image-to-Video Transfer Learning

2024-01-01 · CVPR 2024 1 · Wei zhang, Chaoqun Wan, Tongliang Liu, Xinmei Tian 외

Extending large image-text pre-trained models (e.g. CLIP) for video understanding has made significant advancements. To enable the capability of CLIP to perceive dynamic information in videos existing works are dedic…

Transfer LearningVideo Understanding

CLIP4Caption: CLIP for Video Caption

2021-10-13 · Mingkang Tang, Zhanyu Wang, Zhenhua Liu, Fengyun Rao 외

Video captioning is a challenging task since it requires generating sentences describing various diverse and complex videos. Existing video captioning models lack adequate visual representation due to the neglect of the …

DecoderSentenceText GenerationText Matching+2

Understanding-Enhanced Model Collaboration for Long-Tailed Egocentric Mistake Detection

2026-06-01 · Boyu Han, Qianqian Xu, Shilong Bao, Zhiyong Yang 외 arxiv

In this report, we address the problem of determining whether a user performs an action incorrectly from egocentric video data. To this end, we propose an Understanding-Enhanced Model Collaboration Method (UE-MCM) that c…

From Understanding to Engagement: Personalized pharmacy Video Clips via Vision Language Models (VLMs)

2026-01-08 · Suyash Mishra, Qiang Li, Srikanth Patil, Anubhav Girdhar arxiv

Vision Language Models (VLMs) are poised to revolutionize the digital transformation of pharmacyceutical industry by enabling intelligent, scalable, and automated multi-modality content processing. Traditional manual ann…

Video Summarization

SVAC: Scaling Is All You Need For Referring Video Object Segmentation

2025-09-28 · Li Zhang, Haoxiang Gao, Zhihao Zhang, Luoxiao Huang 외 arxiv

Referring Video Object Segmentation (RVOS) aims to segment target objects in video sequences based on natural language descriptions. While recent advances in Multi-modal Large Language Models (MLLMs) have improved RVOS p…

Referring Video Object Segmentation