paper-with-me

Papers

OmniTraffic: A Controllable Generation Pipeline and Benchmark for Spatio-Temporal Traffic Reasoning

2026-06-14 · Maonan Wang, Zhengyan Huang, Kemou Jiang, Yuhang Fu, Jiayue Zhu, Yuxin Cai, Xingchen Zou, Qiaosheng Zhang, Yi Yu, Ding Wang, Xi Chen, Ben M. Chen, Yuxuan Liang, Zhiyong Cui, Man On Pun, Yirong Chen arxiv

Traffic scene understanding requires models to reason beyond object recognition, including lane topology, multi-view geometry, temporal evolution, and signal-phase semantics. However, existing traffic-oriented multimodal benchmarks largely emphasize passive visual recognition or isolated video understanding, offering limited support for evaluating structure-aware traffic reasoning under controlled conditions. We introduce OmniTraffic, a controllable generation pipeline and benchmark for spatio-temporal traffic reasoning. Built around 12 real-world intersections reconstructed into editable 3D traffic environments and complemented by surveillance footage from two countries, OmniTraffic supports both controlled and natural-condition evaluation. It defines a three-level task hierarchy spanning scene perception, multi-view and temporal reasoning, and decision support. Using structured traffic metadata, OmniTraffic generates synchronized multi-view VQA samples covering vehicle states, lane functions, view--BEV correspondence, temporal dynamics, and signal-phase analysis, resulting in 8M VQA samples and a 3K human-verified test set. Evaluation of eleven frontier MLLMs reveals a large human--model gap, with the most pronounced failures in topology-grounded and spatio-temporal reasoning tasks. Fine-tuning a lightweight MLLM on simulated OmniTraffic data further improves performance on real-world traffic scenes, demonstrating the value of simulation-generated supervision for traffic-specific multimodal reasoning. Beyond a fixed dataset, OmniTraffic provides an extensible pipeline with configurable intersections, camera views, traffic demands, signal phases, visual conditions, and rare events.

📄 PDF Abstract BibTeX arXiv:2606.15749

Code (0)

등록된 구현이 없습니다.

Tasks

Multimodal ReasoningScene UnderstandingObject Recognition

Similar Papers 제목 키워드 기반

AniSora: Exploring the Frontiers of Animation Video Generation in the Sora Era

2024-12-13 · Yudong Jiang, Baohan Xu, Siqian Yang, Mingyu Yin 외

Animation has gained significant interest in the recent film and TV industry. Despite the success of advanced video generation models like Sora, Kling, and CogVideoX in generating natural videos, they lack the same effec…

Image to Video GenerationVideo Generation

MultiShotMaster: A Controllable Multi-Shot Video Generation Framework

2025-12-02 · Qinghe Wang, Xiaoyu Shi, Baolu Li, Weikang Bian 외 arxiv

Current video generation techniques excel at single-shot clips but struggle to produce narrative multi-shot videos, which require flexible shot arrangement, coherent narrative, and controllability beyond text prompts. To…

Video Generation

InSpatio-WorldFM: An Open-Source Real-Time Generative Frame Model

2026-03-12 · InSpatio Team, Donghui Shen, Guofeng Zhang, Haomin Liu 외 arxiv

We present InSpatio-WorldFM, an open-source real-time frame model for spatial intelligence. Unlike video-based world models that rely on sequential frame generation and incur substantial latency due to window-level proce…

InfiniVerse: Occupancy Guided Unbounded Scene Generation for Autonomous Driving

2026-06-30 · Xiaoyu Ye, Leheng Li, Xinyu Ji, Yingjie Cai 외 arxiv

Generating realistic, controllable, and temporally coherent urban environments is a critical yet unresolved challenge in the autonomous driving community. In this paper, we introduce InfiniVerse, a unified pipeline for l…

Autonomous DrivingScene Generation

INSPATIO-WORLD: A Real-Time 4D World Simulator via Spatiotemporal Autoregressive Modeling

2026-04-08 · InSpatio Team, Donghui Shen, Guofeng Zhang, Haomin Liu 외 arxiv

Building world models with spatial consistency and real-time interactivity remains a fundamental challenge in computer vision. Current video generation paradigms often struggle with a lack of spatial persistence and insu…

Video Generation