paper-with-me

홈 › Papers

SceneScribe-1M: A Large-Scale Video Dataset with Comprehensive Geometric and Semantic Annotations

2026-04-09 · Yunnan Wang, Kecheng Zheng, Jianyuan Wang, Minghao Chen, David Novotny, Christian Rupprecht, Yinghao Xu, Xing Zhu, Wenjun Zeng, Xin Jin, Yujun Shen arxiv

The convergence of 3D geometric perception and video synthesis has created an unprecedented demand for large-scale video data that is rich in both semantic and spatio-temporal information. While existing datasets have advanced either 3D understanding or video generation, a significant gap remains in providing a unified resource that supports both domains at scale. To bridge this chasm, we introduce SceneScribe-1M, a new large-scale, multi-modal video dataset. It comprises one million in-the-wild videos, each meticulously annotated with detailed textual descriptions, precise camera parameters, dense depth maps, and consistent 3D point tracks. We demonstrate the versatility and value of SceneScribe-1M by establishing benchmarks across a wide array of downstream tasks, including monocular depth estimation, scene reconstruction, and dynamic point tracking, as well as generative tasks such as text-to-video synthesis, with or without camera control. By open-sourcing SceneScribe-1M, we aim to provide a comprehensive benchmark and a catalyst for research, fostering the development of models that can both perceive the dynamic 3D world and generate controllable, realistic video content.

📄 PDF Abstract BibTeX arXiv:2604.07990

Code (0)

등록된 구현이 없습니다.

Tasks

Monocular Depth EstimationVideo GenerationPoint Tracking

Similar Papers 제목 키워드 기반

Celeb-DF: A Large-scale Challenging Dataset for DeepFake Forensics

2019-09-27 · CVPR 2020 6 · Yuezun Li, Xin Yang, Pu Sun, Honggang Qi 외

AI-synthesized face-swapping videos, commonly known as DeepFakes, is an emerging problem threatening the trustworthiness of online information. The need to develop and evaluate DeepFake detection algorithms calls for lar…

DeepFake DetectionFace Swapping

Actionet: An Interactive End-To-End Platform For Task-Based Data Collection And Augmentation In 3D Environment

2020-10-03 · Jiafei Duan, Samson Yu, Hui Li Tan, Cheston Tan

The problem of task planning for artificial agents remains largely unsolved. While there has been increasing interest in data-driven approaches for the study of task planning for artificial agents, a significant remainin…

Dataset GenerationTask Planning

FriendsQA: A New Large-Scale Deep Video Understanding Dataset with Fine-grained Topic Categorization for Story Videos

2024-12-22 · Zhengqian Wu, Ruizhe Li, Zijun Xu, Zhongyuan Wang 외

Video question answering (VideoQA) aims to answer natural language questions according to the given videos. Although existing models perform well in the factoid VideoQA task, they still face challenges in deep video unde…

Language ModellingLarge Language ModelQuestion AnsweringVideo Question Answering+1

A Large-scale Comprehensive Dataset and Copy-overlap Aware Evaluation Protocol for Segment-level Video Copy Detection

2022-03-05 · CVPR 2022 1 · Sifeng He, Xudong Yang, Chen Jiang, Gang Liang 외

In this paper, we introduce VCSL (Video Copy Segment Localization), a new comprehensive segment-level annotated video copy dataset. Compared with existing copy detection datasets restricted by either video-level annotati…

BenchmarkingCopy Detection

FVQ: A Large-Scale Dataset and A LMM-based Method for Face Video Quality Assessment

2025-04-12 · Sijing Wu, Yunhao Li, Ziwen Xu, Yixuan Gao 외

Face video quality assessment (FVQA) deserves to be explored in addition to general video quality assessment (VQA), as face videos are the primary content on social media platforms and human visual system (HVS) is partic…

Video Quality AssessmentVisual Question Answering (VQA)