paper-with-me

홈 › Papers

LSA: Localized Semantic Alignment for Enhancing Temporal Consistency in Traffic Video Generation

2026-02-05 · Mirlan Karimov, Teodora Spasojevic, Markus Braun, Julian Wiederer, Vasileios Belagiannis, Marc Pollefeys arxiv

Controllable video generation has emerged as a versatile tool for autonomous driving, enabling realistic synthesis of traffic scenarios. However, existing methods depend on control signals at inference time to guide the generative model towards temporally consistent generation of dynamic objects, limiting their utility as scalable and generalizable data engines. In this work, we propose Localized Semantic Alignment (LSA), a simple yet effective framework for fine-tuning pre-trained video generation models. LSA enhances temporal consistency by aligning semantic features between ground-truth and generated video clips. Specifically, we compare the output of an off-the-shelf feature extraction model between the ground-truth and generated video clips localized around dynamic objects inducing a semantic feature consistency loss. We fine-tune the base model by combining this loss with the standard diffusion loss. The model fine-tuned for a single epoch with our novel loss outperforms the baselines in common video generation evaluation metrics. To further test the temporal consistency in generated videos we adapt two additional metrics from object detection task, namely mAP and mIoU. Extensive experiments on nuScenes and KITTI datasets show the effectiveness of our approach in enhancing temporal consistency in video generation without the need for external control signals during inference and any computational overheads.

📄 PDF Abstract BibTeX arXiv:2602.05966

Code (0)

등록된 구현이 없습니다.

Tasks

Autonomous DrivingObject DetectionVideo Generation

Similar Papers 제목 키워드 기반

Enhancing LLMs for Time Series Forecasting via Structure-Guided Cross-Modal Alignment

2025-05-19 · Siming Sun, Kai Zhang, Xuejun Jiang, Wenchao Meng 외

The emerging paradigm of leveraging pretrained large language models (LLMs) for time series forecasting has predominantly employed linguistic-temporal modality alignment strategies through token-level or layer-wise featu…

cross-modal alignmentTime SeriesTime Series Forecasting

MASRA: MLLM-Assisted Semantic-Relational Consistent Alignment for Video Temporal Grounding

2026-05-05 · Ran Ran, Jiwei Wei, Shuchang Zhou, Yitong Qin 외 arxiv

Video Temporal Grounding (VTG) faces a cross-modal semantic gap that often leads to background features being incorrectly aligned with the query, while directly matching the query to moments results in insufficient discr…

Enhancing Low-Cost Video Editing with Lightweight Adaptors and Temporal-Aware Inversion

2025-01-08 · Yangfan He, Sida Li, Kun Li, Xinyuan Song 외

Recent advancements in text-to-image (T2I) generation using diffusion models have enabled cost-effective video-editing applications by leveraging pre-trained models, eliminating the need for resource-intensive training. …

Video Editing

FlowDirector: Training-Free Flow Steering for Precise Text-to-Video Editing

2025-06-05 · Guangzhao Li, Yanming Yang, Chenxi Song, Chi Zhang

Text-driven video editing aims to modify video content according to natural language instructions. While recent training-free approaches have made progress by leveraging pre-trained diffusion models, they typically rely …

Text-to-Video EditingVideo Editing

PredNext: Explicit Cross-View Temporal Prediction for Unsupervised Learning in Spiking Neural Networks

2025-09-29 · Yiting Dong, Jianhao Ding, Zijie Xu, Tong Bu 외 arxiv

Spiking Neural Networks (SNNs), with their temporal processing capabilities and biologically plausible dynamics, offer a natural platform for unsupervised representation learning. However, current unsupervised SNNs predo…

Self-Supervised LearningRepresentation Learning