paper-with-me

홈 › Papers

Trajectory Forcing: Structure-First Generation with Controllable Semantic Trajectories

2026-06-21 · Merve Kocabas, Gege Gao, Bernhard Schölkopf, Andreas Geiger arxiv

Diffusion and flow-based generative models produce strong images, yet their controllability remains largely endpoint-centric: users specify conditions and receive final outputs, while the intermediate generative dynamics remain hidden. Recent methods have begun to exploit generation order and process decomposition to improve sample quality, but still treat intermediate states as internal computation rather than objects for interaction. We propose Trajectory Forcing (TF), a trajectory-centric framework that makes the generation path explicit, semantic, and editable. TF organizes synthesis as a sequence of semantically structured stages, progressing from global layout to object-, part-, and detail-level representations. Each stage produces a decodable latent state that can be inspected, evaluated, and locally edited before the next stage begins. To instantiate this path, we derive coarse-to-fine teacher hierarchies by clustering pretrained visual representations such as DINOv2, and train a hierarchy-conditioned one-step flow-matching model at each level. We further introduce trajectory-aware metrics that measure structural consistency and local controllability beyond endpoint quality metrics such as FID. Experiments show that TF achieves competitive sample quality while exposing coherent intermediate states and supporting localized edits across semantic levels. By shifting the focus from final images to the generative path itself, TF opens a route toward controllable, trajectory-aware image synthesis.

📄 PDF Abstract BibTeX arXiv:2606.22527

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

FlashMotion: Few-Step Controllable Video Generation with Trajectory Guidance

2026-03-12 · Quanhao Li, Zhen Xing, Rui Wang, Haidong Cao 외 arxiv

Recent advances in trajectory-controllable video generation have achieved remarkable progress. Previous methods mainly use adapter-based architectures for precise motion control along predefined trajectories. However, al…

Video Generation

SentBS: Sentence-level Beam Search for Controllable Summarization

2022-10-26 · Chenhui Shen, Liying Cheng, Lidong Bing, Yang You 외

A wide range of control perspectives have been explored in controllable text generation. Structure-controlled summarization is recently proposed as a useful and interesting research direction. However, current structure-…

SentenceText Generation

Generative Active Learning for Long-tail Trajectory Prediction via Controllable Diffusion Model

2025-07-30 · Daehee Park, Monu Surana, Pranav Desai, Ashish Mehta 외 arxiv

While data-driven trajectory prediction has enhanced the reliability of autonomous driving systems, it still struggles with rarely observed long-tail scenarios. Prior works addressed this by modifying model architectures…

Trajectory PredictionLong-tail LearningAutonomous DrivingActive Learning

FreeTraj: Tuning-Free Trajectory Control in Video Diffusion Models

2024-06-24 · Haonan Qiu, Zhaoxi Chen, Zhouxia Wang, Yingqing He 외

Diffusion model has demonstrated remarkable capability in video generation, which further sparks interest in introducing trajectory control into the generation process. While existing works mainly focus on training-based…

Video Generation

minWM: A Full-Stack Open-Source Framework for Real-Time Interactive Video World Models

2026-05-28 · Min Zhao, Hongzhou Zhu, Bokai Yan, Zihan Zhou 외 arxiv

Recent video diffusion foundation models have achieved remarkable progress in high-quality video generation, yet turning them into real-time interactive video world models remains challenging. Interactive world models re…

Video Generation