paper-with-me

홈 › Papers

Watch and Learn: Mapping Language and Noisy Real-world Videos with Self-supervision

2020-11-19 · Yujie Zhong, Linhai Xie, Sen Wang, Lucia Specia, Yishu Miao

In this paper, we teach machines to understand visuals and natural language by learning the mapping between sentences and noisy video snippets without explicit annotations. Firstly, we define a self-supervised learning framework that captures the cross-modal information. A novel adversarial learning module is then introduced to explicitly handle the noises in the natural videos, where the subtitle sentences are not guaranteed to be strongly corresponded to the video snippets. For training and evaluation, we contribute a new dataset `ApartmenTour' that contains a large number of online videos and subtitles. We carry out experiments on the bidirectional retrieval tasks between sentences and videos, and the results demonstrate that our proposed model achieves the state-of-the-art performance on both retrieval tasks and exceeds several strong baselines. The dataset can be downloaded at https://github.com/zyj-13/WAL.

📄 PDF Abstract BibTeX arXiv:2011.09634

Code (1)

zyj-13/WAL 공식 구현

Tasks

RetrievalSelf-Supervised Learning

Similar Papers 제목 키워드 기반

Uncovering User Interest from Biased and Noised Watch Time in Video Recommendation

2023-08-16 · Haiyuan Zhao, Lei Zhang, Jun Xu, Guohao Cai 외

In the video recommendation, watch time is commonly adopted as an indicator of user interest. However, watch time is not only influenced by the matching of users' interests but also by other factors, such as duration bia…

Mapping Between fMRI Responses to Movies and their Natural Language Annotations

2016-10-13 · Kiran Vodrahalli, Po-Hsuan Chen, YIngyu Liang, Christopher Baldassano 외

Several research groups have shown how to correlate fMRI responses to the meanings of presented stimuli. This paper presents new methods for doing so when only a natural language annotation is available as the descriptio…

Scene ClassificationSentenceSentence EmbeddingSentence-Embedding

BayesBeat: Reliable Atrial Fibrillation Detection from Noisy Photoplethysmography Data

2020-11-02 · Sarkar Snigdha Sarathi Das, Subangkar Karmaker Shanto, Masum Rahman, Md. Saiful Islam 외

Smartwatches or fitness trackers have garnered a lot of popularity as potential health tracking devices due to their affordable and longitudinal monitoring capabilities. To further widen their health tracking capabilitie…

Atrial Fibrillation DetectionPhotoplethysmography (PPG)

SafeWatch: An Efficient Safety-Policy Following Video Guardrail Model with Transparent Explanations

2024-12-09 · Zhaorun Chen, Francesco Pinto, Minzhou Pan, Bo Li

With the rise of generative AI and rapid growth of high-quality video generation, video guardrails have become more crucial than ever to ensure safety and security across platforms. Current video guardrails, however, are…

Video Generation

Detecting In-Person Conversations in Noisy Real-World Environments with Smartwatch Audio and Motion Sensing

2025-07-16 · Alice Zhang, Callihan Bertley, Dawei Liang, Edison Thomaz arxiv

Social interactions play a crucial role in shaping human behavior, relationships, and societies. It encompasses various forms of communication, such as verbal conversation, non-verbal gestures, facial expressions, and bo…