paper-with-me

Papers

ProTA: Probabilistic Token Aggregation for Text-Video Retrieval

2024-04-18 · Han Fang, Xianghao Zang, Chao Ban, Zerun Feng, Lanxiang Zhou, Zhongjiang He, Yongxiang Li, Hao Sun

Text-video retrieval aims to find the most relevant cross-modal samples for a given query. Recent methods focus on modeling the whole spatial-temporal relations. However, since video clips contain more diverse content than captions, the model aligning these asymmetric video-text pairs has a high risk of retrieving many false positive results. In this paper, we propose Probabilistic Token Aggregation (ProTA) to handle cross-modal interaction with content asymmetry. Specifically, we propose dual partial-related aggregation to disentangle and re-aggregate token representations in both low-dimension and high-dimension spaces. We propose token-based probabilistic alignment to generate token-level probabilistic representation and maintain the feature representation diversity. In addition, an adaptive contrastive loss is proposed to learn compact cross-modal distribution space. Based on extensive experiments, ProTA achieves significant improvements on MSR-VTT (50.9%), LSMDC (25.8%), and DiDeMo (47.2%).

📄 PDF Abstract BibTeX arXiv:2404.12216

Code (0)

등록된 구현이 없습니다.

Tasks

DiversityRetrievalVideo Retrieval

Methods 이 논문이 사용한 방법론

Focus 설명 없음

Similar Papers 제목 키워드 기반

Make-A-Protagonist: Generic Video Editing with An Ensemble of Experts

2023-05-15 · Yuyang Zhao, Enze Xie, Lanqing Hong, Zhenguo Li 외

The text-driven image and video diffusion models have achieved unprecedented success in generating realistic and diverse content. Recently, the editing and variation of existing images and videos in diffusion-based gener…

DenoisingVideo EditingVideo Generation

ProTAL: A Drag-and-Link Video Programming Framework for Temporal Action Localization

2025-05-23 · Yuchen He, Jianbing Lv, Liqi Cheng, Lingyu Meng 외

Temporal Action Localization (TAL) aims to detect the start and end timestamps of actions in a video. However, the training of TAL models requires a substantial amount of manually annotated data. Data programming is an e…

Action LocalizationTemporal Action Localization

SAVE: Protagonist Diversification with Structure Agnostic Video Editing

2023-12-05 · Yeji Song, Wonsik Shin, Junsoo Lee, Jeesoo Kim 외

Driven by the upsurge progress in text-to-image (T2I) generation models, text-to-video (T2V) generation has experienced a significant advance as well. Accordingly, tasks such as modifying the object or changing the style…

Optical Flow EstimationVideo Editing

TESTA: Temporal-Spatial Token Aggregation for Long-form Video-Language Understanding

2023-10-29 · Shuhuai Ren, Sishuo Chen, Shicheng Li, Xu sun 외

Large-scale video-language pre-training has made remarkable strides in advancing video-language understanding tasks. However, the heavy computational burden of video encoding remains a formidable efficiency bottleneck, p…

FormLanguage ModellingRetrievalVideo Question Answering+2

Beyond Frame Selection: Generative Latent Evidence Aggregation for Long-Video Understanding

2026-07-30 · Bowen Liu, Shuning Wang, Xinpeng Ding, Zhiheng Wu 외 arxiv

Long-video understanding commonly compresses videos into a small set of frames or visual tokens for answer generation. Existing compact pipelines focus on retaining relevant visual content as explicit evidence. Yet makin…

Answer Generation