paper-with-me

Papers

Look at What I'm Doing: Self-Supervised Spatial Grounding of Narrations in Instructional Videos

2021-10-20 · Reuben Tan, Bryan A. Plummer, Kate Saenko, Hailin Jin, Bryan Russell

We introduce the task of spatially localizing narrated interactions in videos. Key to our approach is the ability to learn to spatially localize interactions with self-supervision on a large corpus of videos with accompanying transcribed narrations. To achieve this goal, we propose a multilayer cross-modal attention network that enables effective optimization of a contrastive loss during training. We introduce a divided strategy that alternates between computing inter- and intra-modal attention across the visual and natural language modalities, which allows effective training via directly contrasting the two modalities' representations. We demonstrate the effectiveness of our approach by self-training on the HowTo100M instructional video dataset and evaluating on a newly collected dataset of localized described interactions in the YouCook2 dataset. We show that our approach outperforms alternative baselines, including shallow co-attention and full cross-modal attention. We also apply our approach to grounding phrases in images with weak supervision on Flickr30K and show that stacking multiple attention layers is effective and, when combined with a word-to-region loss, achieves state of the art on recall-at-one and pointing hand accuracies.

📄 PDF Abstract BibTeX arXiv:2110.10596

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Look at What I’m Doing: Self-Supervised Spatial Grounding of Narrations in Instructional Videos

2021-12-01 · NeurIPS 2021 12 · Reuben Tan, Bryan Plummer, Kate Saenko, Hailin Jin 외

We introduce the task of spatially localizing narrated interactions in videos. Key to our approach is the ability to learn to spatially localize interactions with self-supervision on a large corpus of videos with accompa…

Self-supervised Monocular Depth Estimation with Large Kernel Attention

2024-09-26 · Xuezhi Xiang, Yao Wang, Lei Zhang, Denis Ombati 외

Self-supervised monocular depth estimation has emerged as a promising approach since it does not rely on labeled training data. Most methods combine convolution and Transformer to model long-distance dependencies to esti…

DecoderDepth EstimationMonocular Depth Estimation

LLMs and the ZPD

2026-05-12 · Peter Wallis arxiv

One hundred years ago Vygotsky and his circle were exploring the nature of consciousness and defining what would become psychology in the Soviet Union. They concluded that children develop "scientific thinking" through i…

Unsupervised Object Localization: Observing the Background to Discover Objects

2022-12-15 · CVPR 2023 1 · Oriane Siméoni, Chloé Sekkat, Gilles Puy, Antonin Vobecky 외

Recent advances in self-supervised visual representation learning have paved the way for unsupervised methods tackling tasks such as object discovery and instance segmentation. However, discovering objects in an image wi…

Instance SegmentationObjectObject DiscoveryObject Localization+7

Self-Supervised Learning of Face Representations for Video Face Clustering

2019-03-03 · Vivek Sharma, Makarand Tapaswi, M. Saquib Sarfraz, Rainer Stiefelhagen

Analyzing the story behind TV series and movies often requires understanding who the characters are and what they are doing. With improving deep face models, this may seem like a solved problem. However, as face detector…

ClusteringDiversityFace ClusteringSelf-Supervised Learning