Motion-Based Weak Supervision for Video Parsing with Application to Colonoscopy
We propose a two-stage unsupervised approach for parsing videos into phases. We use motion cues to divide the video into coarse segments. Noisy segment labels are then used to weakly supervise an appearance-based classifier. We show the effectiveness of the method for phase detection in colonoscopy videos.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
FlowVVTON: Flow-Guided Mask-Free Video Virtual Try-On
Video virtual try-on aims to transfer a target garment onto a moving person across video frames. Current methods rely on human parsing masks or pose keypoints that frequently fail under large motions and occlusions, caus…
Virtual Try-onHuman ParsingTeacher-Guided Pseudo Supervision and Cross-Modal Alignment for Audio-Visual Video Parsing
Weakly-supervised audio-visual video parsing (AVVP) seeks to detect audible, visible, and audio-visual events without temporal annotations. Previous work has emphasized refining global predictions through contrastive or …
DocParser: Hierarchical Structure Parsing of Document Renderings
Translating renderings (e. g. PDFs, scans) into hierarchical document structures is extensively demanded in the daily routines of many real-world applications. However, a holistic, principled approach to inferring the co…
Weakly-Supervised Audio-Visual Video Parsing with Prototype-based Pseudo-Labeling
In this paper we address the weakly-supervised Audio-Visual Video Parsing (AVVP) problem which aims at labeling events in a video as audible visible or both and temporally localizing and classifying them into known c…
Contrastive LearningMultiple Instance LearningAdvancing Weakly-Supervised Audio-Visual Video Parsing via Segment-wise Pseudo Labeling
The Audio-Visual Video Parsing task aims to identify and temporally localize the events that occur in either or both the audio and visual streams of audible videos. It often performs in a weakly-supervised manner, where …
audio-visual event localizationDenoisingPseudo Label