paper-with-me

홈 › Papers

CAST: Character labeling in Animation using Self-supervision by Tracking

2022-01-19 · Oron Nir, Gal Rapoport, Ariel Shamir

Cartoons and animation domain videos have very different characteristics compared to real-life images and videos. In addition, this domain carries a large variability in styles. Current computer vision and deep-learning solutions often fail on animated content because they were trained on natural images. In this paper we present a method to refine a semantic representation suitable for specific animated content. We first train a neural network on a large-scale set of animation videos and use the mapping to deep features as an embedding space. Next, we use self-supervision to refine the representation for any specific animation style by gathering many examples of animated characters in this style, using a multi-object tracking. These examples are used to define triplets for contrastive loss training. The refined semantic space allows better clustering of animated characters even when they have diverse manifestations. Using this space we can build dictionaries of characters in an animation videos, and define specialized classifiers for specific stylistic content (e.g., characters in a specific animation series) with very little user effort. These classifiers are the basis for automatically labeling characters in animation videos. We present results on a collection of characters in a variety of animation styles.

📄 PDF Abstract BibTeX arXiv:2201.07619

Code (1)

oronnir/cast 공식 구현 pytorch

Tasks

Multi-Object TrackingObject TrackingRepresentation Learning

Methods 이 논문이 사용한 방법론

Supervised Contrastive Loss 설명 없음

Similar Papers 제목 키워드 기반

STARCaster: Spatio-Temporal AutoRegressive Video Diffusion for Identity- and View-Aware Talking Portraits

2025-12-15 · Foivos Paraperas Papantoniou, Stathis Galanakis, Rolandos Alexandros Potamias, Bernhard Kainz 외 arxiv

This paper presents STARCaster, an identity-aware spatio-temporal video diffusion model that addresses both speech-driven portrait animation and free-viewpoint talking portrait synthesis, given an identity embedding or r…

Lip Reading

MarioNette: Self-Supervised Sprite Learning

2021-04-29 · NeurIPS 2021 12 · Dmitriy Smirnov, Michael Gharbi, Matthew Fisher, Vitor Guizilini 외

Artists and video game designers often construct 2D animations using libraries of sprites -- textured patches of objects and characters. We propose a deep learning approach that decomposes sprite-based video animations i…

PerformRecast: Expression and Head Pose Disentanglement for Portrait Video Editing

2026-03-20 · Jiadong Liang, Bojun Xiong, Jie Tian, Hua Li 외 arxiv

This paper primarily investigates the task of expression-only portrait video performance editing based on a driving video, which plays a crucial role in animation and film industries. Most existing research mainly focuse…

Automatic segmentation of meniscus based on MAE self-supervision and point-line weak supervision paradigm

2022-05-07 · Yuhan Xie, Kexin Jiang, Zhiyong Zhang, Shaolong Chen 외

Medical image segmentation based on deep learning is often faced with the problems of insufficient datasets and long time-consuming labeling. In this paper, we introduce the self-supervised method MAE(Masked Autoencoders…

Image SegmentationMedical Image SegmentationPseudo LabelSegmentation+1

Generalized Weak Supervision for Neural Information Retrieval

2023-04-18 · Yen-Chieh Lien, Hamed Zamani, W. Bruce Croft

Neural ranking models (NRMs) have demonstrated effective performance in several information retrieval (IR) tasks. However, training NRMs often requires large-scale training data, which is difficult and expensive to obtai…

Information RetrievalPassage RetrievalRetrieval