paper-with-me

홈 › Papers

Distilled Semantics for Comprehensive Scene Understanding from Videos

2020-03-31 · CVPR 2020 6 · Fabio Tosi, Filippo Aleotti, Pierluigi Zama Ramirez, Matteo Poggi, Samuele Salti, Luigi Di Stefano, Stefano Mattoccia

Whole understanding of the surroundings is paramount to autonomous systems. Recent works have shown that deep neural networks can learn geometry (depth) and motion (optical flow) from a monocular video without any explicit supervision from ground truth annotations, particularly hard to source for these two tasks. In this paper, we take an additional step toward holistic scene understanding with monocular cameras by learning depth and motion alongside with semantics, with supervision for the latter provided by a pre-trained network distilling proxy ground truth images. We address the three tasks jointly by a) a novel training protocol based on knowledge distillation and self-supervision and b) a compact network architecture which enables efficient scene understanding on both power hungry GPUs and low-power embedded platforms. We thoroughly assess the performance of our framework and show that it yields state-of-the-art results for monocular depth estimation, optical flow and motion segmentation.

📄 PDF Abstract BibTeX arXiv:2003.14030

Code (1)

CVLAB-Unibo/omeganet 공식 구현 tf

Tasks

Depth EstimationKnowledge DistillationMonocular Depth EstimationMotion SegmentationOptical Flow EstimationScene Understanding

Methods 이 논문이 사용한 방법론

Knowledge Distillation A very simple way to improve the performance of almost any machine learning algorithm is to train many different models on the same data and then to average their predictions.…

Similar Papers 제목 키워드 기반

Multimodal High-order Relation Transformer for Scene Boundary Detection

2023-01-01 · ICCV 2023 1 · Xi Wei, Zhangxiang Shi, Tianzhu Zhang, Xiaoyuan Yu 외

Scene boundary detection breaks down long videos into meaningful story-telling units and plays a crucial role in high-level video understanding. Despite significant advancements in this area, this task remains a chal…

Boundary DetectionDecoderRelationVideo Understanding

Semantic Lens: Instance-Centric Semantic Alignment for Video Super-Resolution

2023-12-13 · Qi Tang, Yao Zhao, Meiqin Liu, Jian Jin 외

As a critical clue of video super-resolution (VSR), inter-frame alignment significantly impacts overall performance. However, accurate pixel-level alignment is a challenging task due to the intricate motion interweaving …

Super-ResolutionVideo Super-Resolution

DyST: Towards Dynamic Neural Scene Representations on Real-World Videos

2023-10-09 · Maximilian Seitzer, Sjoerd van Steenkiste, Thomas Kipf, Klaus Greff 외

Visual understanding of the world goes beyond the semantics and flat structure of individual images. In this work, we aim to capture both the 3D structure and dynamics of real-world scenes from monocular real-world video…

UGC-VideoCaptioner: An Omni UGC Video Detail Caption Model and New Benchmarks

2025-07-15 · Peiran Wu, Yunze Liu, Zhengdong Zhu, Enmin Zhou 외

Real-world user-generated videos, especially on platforms like TikTok, often feature rich and intertwined audio visual content. However, existing video captioning benchmarks and models remain predominantly visual centric…

Video CaptioningVideo Understanding

MSC: A Marine Wildlife Video Dataset with Grounded Segmentation and Clip-Level Captioning

2025-08-06 · Quang-Trung Truong, Yuk-Kwan Wong, Vo Hoang Kim Tuyen Dang, Rinaldi Gotama 외 arxiv

Marine videos present significant challenges for video understanding due to the dynamics of marine objects and the surrounding environment, camera motion, and the complexity of underwater scenes. Existing video captionin…

Video CaptioningVisual GroundingVideo Generation