paper-with-me

홈 › Papers

EgoVid-5M: A Large-Scale Video-Action Dataset for Egocentric Video Generation

2024-11-13 · XiaoFeng Wang, Kang Zhao, Feng Liu, Jiayu Wang, Guosheng Zhao, Xiaoyi Bao, Zheng Zhu, Yingya Zhang, Xingang Wang

Video generation has emerged as a promising tool for world simulation, leveraging visual data to replicate real-world environments. Within this context, egocentric video generation, which centers on the human perspective, holds significant potential for enhancing applications in virtual reality, augmented reality, and gaming. However, the generation of egocentric videos presents substantial challenges due to the dynamic nature of egocentric viewpoints, the intricate diversity of actions, and the complex variety of scenes encountered. Existing datasets are inadequate for addressing these challenges effectively. To bridge this gap, we present EgoVid-5M, the first high-quality dataset specifically curated for egocentric video generation. EgoVid-5M encompasses 5 million egocentric video clips and is enriched with detailed action annotations, including fine-grained kinematic control and high-level textual descriptions. To ensure the integrity and usability of the dataset, we implement a sophisticated data cleaning pipeline designed to maintain frame consistency, action coherence, and motion smoothness under egocentric conditions. Furthermore, we introduce EgoDreamer, which is capable of generating egocentric videos driven simultaneously by action descriptions and kinematic control signals. The EgoVid-5M dataset, associated action annotations, and all data cleansing metadata will be released for the advancement of research in egocentric video generation.

📄 PDF Abstract BibTeX arXiv:2411.08380

Code (0)

등록된 구현이 없습니다.

Tasks

Video Generation

Similar Papers 제목 키워드 기반

EgoVideo: Exploring Egocentric Foundation Model and Downstream Adaptation

2024-06-26 · Baoqi Pei, Guo Chen, Jilan Xu, Yuping He 외

In this report, we present our solutions to the EgoVis Challenges in CVPR 2024, including five tracks in the Ego4D challenge and three tracks in the EPIC-Kitchens challenge. Building upon the video-language two-tower mod…

Action AnticipationAction RecognitionDomain AdaptationLong Term Action Anticipation+4

Modeling Fine-Grained Hand-Object Dynamics for Egocentric Video Representation Learning

2025-03-02 · Baoqi Pei, Yifei HUANG, Jilan Xu, Guo Chen 외

In egocentric video understanding, the motion of hands and objects as well as their interactions play a significant role by nature. However, existing egocentric video representation learning methods mainly focus on align…

Large Language ModelMulti-Instance RetrievalObjectRepresentation Learning+2

An Egocentric Vision-Language Model based Portable Real-time Smart Assistant

2025-03-06 · Yifei HUANG, Jilan Xu, Baoqi Pei, Yuping He 외

We present Vinci, a vision-language system designed to provide real-time, comprehensive AI assistance on portable devices. At its core, Vinci leverages EgoVideo-VL, a novel model that integrates an egocentric vision foun…

Language ModelingLanguage ModellingLarge Language ModelScene Understanding+1

Forensic Video Steganalysis in Spatial Domain by Noise Residual Convolutional Neural Network

2023-05-29 · Mart Keizer, Zeno Geradts, Meike Kombrink

This research evaluates a convolutional neural network (CNN) based approach to forensic video steganalysis. A video steganography dataset is created to train a CNN to conduct forensic steganalysis in the spatial domain. …

Steganalysis

Motion-Focused Latent Action Enables Cross-Embodiment VLA Training from Human EgoVideos

2026-06-17 · Runze Xu, Yiluo Zhang, Jian Wang, Yu Wang 외 arxiv

Training generalist Vision-Language-Action(VLA) models typically requires massive, diverse robotic datasets with high-fidelity action annotations. While egocentric human manipulation videos are abundant and capture signi…