paper-with-me

홈 › Papers

From Seeing to Predicting: A Vision-Language Framework for Trajectory Forecasting and Controlled Video Generation

2025-10-01 · Fan Yang, Zhiyang Chen, Yousong Zhu, Xin Li, Jinqiao Wang arxiv

Current video generation models produce physically inconsistent motion that violates real-world dynamics. We propose TrajVLM-Gen, a two-stage framework for physics-aware image-to-video generation. First, we employ a Vision Language Model to predict coarse-grained motion trajectories that maintain consistency with real-world physics. Second, these trajectories guide video generation through attention-based mechanisms for fine-grained motion refinement. We build a trajectory prediction dataset based on video tracking data with realistic motion patterns. Experiments on UCF-101 and MSR-VTT demonstrate that TrajVLM-Gen outperforms existing methods, achieving competitive FVD scores of 545 on UCF-101 and 539 on MSR-VTT.

📄 PDF Abstract BibTeX arXiv:2510.00806

Code (0)

등록된 구현이 없습니다.

Tasks

Trajectory ForecastingTrajectory PredictionVideo Generation

Similar Papers 제목 키워드 기반

Seeing is Believing (and Predicting): Context-Aware Multi-Human Behavior Prediction with Vision Language Models

2025-12-17 · Utsav Panchal, Yuchen Liu, Luigi Palmieri, Ilche Georgievski 외 arxiv

Accurately predicting human behaviors is crucial for mobile robots operating in human-populated environments. While prior research primarily focuses on predicting actions in single-human scenarios from an egocentric view…

VLA-Reasoner: Empowering Vision-Language-Action Models with Reasoning via Online Monte Carlo Tree Search

2025-09-26 · Wenkai Guo, Guanxing Lu, Haoyuan Deng, Zhenyu Wu 외 arxiv

Vision-Language-Action models (VLAs) achieve strong performance in general robotic manipulation tasks by scaling imitation learning. However, existing VLAs are limited to predicting short-sighted next-action, which strug…

Density Estimation

The All-Seeing Project: Towards Panoptic Visual Recognition and Understanding of the Open World

2023-08-03 · Weiyun Wang, Min Shi, Qingyun Li, Wenhai Wang 외

We present the All-Seeing (AS) project: a large-scale data and model for recognizing and understanding everything in the open world. Using a scalable data engine that incorporates human feedback and efficient models in t…

AllQuestion AnsweringRetrievalText Retrieval

Seeing without Pixels: Perception from Camera Trajectories

2025-11-26 · Zihui Xue, Kristen Grauman, Dima Damen, Andrew Zisserman 외 arxiv

Can one perceive a video's content without seeing its pixels, just from the camera trajectory-the path it carves through space? This paper is the first to systematically investigate this seemingly implausible question. T…

Camera Pose EstimationContrastive Learning

Beyond Waypoints: A Trajectory-Centric Waypointing Paradigm for Vision-Language Navigation

2026-06-05 · Haoxiang Shi, Xiang Deng, Haoyu Zhang, Qiaohui Chu 외 arxiv

Vision-Language Navigation in Continuous Environments (VLN-CE) requires agents to follow natural-language instructions while navigating in real-world-like environments. Most VLN-CE approach\-es adopt a three-stage framew…

Vision-Language Navigation