paper-with-me

Papers

Semantic Role Aware Correlation Transformer for Text to Video Retrieval

2022-06-26 · Burak Satar, Hongyuan Zhu, Xavier Bresson, Joo Hwee Lim

With the emergence of social media, voluminous video clips are uploaded every day, and retrieving the most relevant visual content with a language query becomes critical. Most approaches aim to learn a joint embedding space for plain textual and visual contents without adequately exploiting their intra-modality structures and inter-modality correlations. This paper proposes a novel transformer that explicitly disentangles the text and video into semantic roles of objects, spatial contexts and temporal contexts with an attention scheme to learn the intra- and inter-role correlations among the three roles to discover discriminative features for matching at different levels. The preliminary results on popular YouCook2 indicate that our approach surpasses a current state-of-the-art method, with a high margin in all metrics. It also overpasses two SOTA methods in terms of two metrics.

📄 PDF Abstract BibTeX arXiv:2206.12849

Code (1)

buraksatar/RoME_video_retrieval 공식 구현 pytorch

Tasks

RetrievalText to Video RetrievalVideo Retrieval

Similar Papers 제목 키워드 기반

RoME: Role-aware Mixture-of-Expert Transformer for Text-to-Video Retrieval

2022-06-26 · Burak Satar, Hongyuan Zhu, Hanwang Zhang, Joo Hwee Lim

Seas of videos are uploaded daily with the popularity of social channels; thus, retrieving the most related video contents with user textual queries plays a more crucial role. Most methods consider only one joint embeddi…

Mixture-of-ExpertsRetrievalText to Video RetrievalVideo Retrieval

ClarET: Pre-training a Correlation-Aware Context-To-Event Transformer for Event-Centric Generation and Classification

2022-03-04 · ACL 2022 5 · Yucheng Zhou, Tao Shen, Xiubo Geng, Guodong Long 외

Generating new events given context with correlated ones plays a crucial role in many event-centric reasoning tasks. Existing works either limit their scope to specific scenarios or overlook event-level correlations. In …

counterfactualFew-Shot Learning

Target-aware Bi-Transformer for Few-shot Segmentation

2023-09-18 · Xianglin Wang, Xiaoliu Luo, Taiping Zhang

Traditional semantic segmentation tasks require a large number of labels and are difficult to identify unlearned categories. Few-shot semantic segmentation (FSS) aims to use limited labeled support images to identify the…

Few-Shot Semantic SegmentationSegmentationSemantic Segmentation

Full Point Encoding for Local Feature Aggregation in 3D Point Clouds

2023-03-08 · Yong He, Hongshan Yu, Zhengeng Yang, Xiaoyan Liu 외

Point cloud processing methods exploit local point features and global context through aggregation which does not explicity model the internal correlations between local and global features. To address this problem, we p…

object-detectionObject DetectionPositionSemantic Segmentation

SA$^{2}$Net: Scale-Adaptive Structure-Affinity Transformation for Spine Segmentation from Ultrasound Volume Projection Imaging

2025-10-30 · Hao Xie, Zixun Huang, Yushen Zuo, Yakun Ju 외 arxiv

Spine segmentation, based on ultrasound volume projection imaging (VPI), plays a vital role for intelligent scoliosis diagnosis in clinical applications. However, this task faces several significant challenges. Firstly, …