paper-with-me

Papers

MovieCLIP: Visual Scene Recognition in Movies

2022-10-20 · Digbalay Bose, Rajat Hebbar, Krishna Somandepalli, Haoyang Zhang, Yin Cui, Kree Cole-McLaughlin, Huisheng Wang, Shrikanth Narayanan

Longform media such as movies have complex narrative structures, with events spanning a rich variety of ambient visual scenes. Domain specific challenges associated with visual scenes in movies include transitions, person coverage, and a wide array of real-life and fictional scenarios. Existing visual scene datasets in movies have limited taxonomies and don't consider the visual scene transition within movie clips. In this work, we address the problem of visual scene recognition in movies by first automatically curating a new and extensive movie-centric taxonomy of 179 scene labels derived from movie scripts and auxiliary web-based video datasets. Instead of manual annotations which can be expensive, we use CLIP to weakly label 1.12 million shots from 32K movie clips based on our proposed taxonomy. We provide baseline visual models trained on the weakly labeled dataset called MovieCLIP and evaluate them on an independent dataset verified by human raters. We show that leveraging features from models pretrained on MovieCLIP benefits downstream tasks such as multi-label scene and genre classification of web videos and movie trailers.

📄 PDF Abstract BibTeX arXiv:2210.11065

Code (1)

usc-sail/mica-MovieCLIP 공식 구현 pytorch

Tasks

Genre classificationScene Recognition

Methods 이 논문이 사용한 방법론

CLIP Contrastive Language-Image Pre-training (CLIP), consisting of a simplified version of ConVIRT trained from scratch, is an efficient method of image representation learning…

Similar Papers 제목 키워드 기반

A Local-to-Global Approach to Multi-modal Movie Scene Segmentation

2020-04-06 · CVPR 2020 6 · Anyi Rao, Linning Xu, Yu Xiong, Guodong Xu 외

Scene, as the crucial unit of storytelling in movies, contains complex activities of actors and their interactions in a physical environment. Identifying the composition of scenes serves as a critical step towards semant…

Action RecognitionScene SegmentationSegmentation

IDA-VLM: Towards Movie Understanding via ID-Aware Large Vision-Language Model

2024-07-10 · Yatai Ji, Shilong Zhang, Jie Wu, Peize Sun 외

The rapid advancement of Large Vision-Language models (LVLMs) has demonstrated a spectrum of emergent capabilities. Nevertheless, current models only focus on the visual content of a single scenario, while their ability …

Language ModelingLanguage ModellingQuestion Answering

Condensed Movies: Story Based Retrieval with Contextual Embeddings

2020-05-08 · Max Bain, Arsha Nagrani, Andrew Brown, Andrew Zisserman

Our objective in this work is long range understanding of the narrative structure of movies. Instead of considering the entire movie, we propose to learn from the `key scenes' of the movie, providing a condensed look at …

RetrievalText to Video RetrievalVideo Retrieval

Character-Centric Understanding of Animated Movies

2025-09-15 · Zhongrui Gui, Junyu Xie, Tengda Han, Weidi Xie 외 arxiv

Animated movies are captivating for their unique character designs and imaginative storytelling, yet they pose significant challenges for existing recognition systems. Unlike the consistent visual patterns detected by co…

Face Recognition

On Negative Sampling for Audio-Visual Contrastive Learning from Movies

2022-04-29 · Mahdi M. Kalayeh, Shervin Ardeshir, Lingyi Liu, Nagendra Kamath 외

The abundance and ease of utilizing sound, along with the fact that auditory clues reveal a plethora of information about what happens in a scene, make the audio-visual space an intuitive choice for representation learni…

Action RecognitionAudio ClassificationContrastive LearningForm+2