paper-with-me

Papers

Action-conditioned video data improves predictability

2024-04-08 · Meenakshi Sarkar, Debasish Ghose

Long-term video generation and prediction remain challenging tasks in computer vision, particularly in partially observable scenarios where cameras are mounted on moving platforms. The interaction between observed image frames and the motion of the recording agent introduces additional complexities. To address these issues, we introduce the Action-Conditioned Video Generation (ACVG) framework, a novel approach that investigates the relationship between actions and generated image frames through a deep dual Generator-Actor architecture. ACVG generates video sequences conditioned on the actions of robots, enabling exploration and analysis of how vision and action mutually influence one another in dynamic environments. We evaluate the framework's effectiveness on an indoor robot motion dataset which consists of sequences of image frames along with the sequences of actions taken by the robotic agent, conducting a comprehensive empirical study comparing ACVG to other state-of-the-art frameworks along with a detailed ablation study.

📄 PDF Abstract BibTeX arXiv:2404.05439

Code (0)

등록된 구현이 없습니다.

Tasks

Video Generation

Similar Papers 제목 키워드 기반

Unified Video-Action Joint Denoising for Dexterous Action and Data Generation

2026-06-02 · Dingrui Wang, YuAn Wang, Jinkun Liu, Yue Zhang 외 arxiv

Recent world action models leverage video foundation models by aligning broad visual-dynamics priors with executable robot actions. We revisit this alignment from a distributional perspective. Existing formulations typic…

Back to the Future: The Role of Past and Future Context Predictability in Incremental Language Production

2025-11-11 · Shiva Upadhye, Richard Futrell arxiv

Contextual predictability shapes how we choose and encode words in production. The effects of a word's predictability given preceding or past context are generally well-understood in both production and comprehension, bu…

RAE-NWM: Navigation World Model in Dense Visual Representation Space

2026-03-10 · Mingkun Zhang, Wangtian Shen, Fan Zhang, Haijian Qin 외 arxiv

Visual navigation requires agents to reach goals in complex environments through perception and planning. World models address this task by simulating action-conditioned state transitions to predict future observations. …

Visual Navigation

GV-VAD : Exploring Video Generation for Weakly-Supervised Video Anomaly Detection

2025-08-01 · Suhang Cai, Xiaohao Peng, Chong Wang, Xiaojie Cai 외 arxiv

Video anomaly detection (VAD) plays a critical role in public safety applications such as intelligent surveillance. However, the rarity, unpredictability, and high annotation cost of real-world anomalies make it difficul…

Weakly-supervised Video Anomaly DetectionVideo Generation

Re-imagining ISO 26262 in the Age of Autonomous Vehicles: Enhancing Controllability through Transferability and Predictability

2026-06-05 · Chaitanya Shinde, Hadi Hajieghrary, Paul Schmitt, Adam Shoemaker 외 arxiv

The ISO 26262 standard defines functional safety for road vehicles through risk assessments based on Severity, Exposure, and Controllability, grounded in a human-driven vehicle paradigm. In the context of autonomous vehi…

Autonomous Vehicles