paper-with-me

Papers

JointDiff: Bridging Continuous and Discrete in Multi-Agent Trajectory Generation

2025-09-26 · Guillem Capellera, Luis Ferraz, Antonio Rubio, Alexandre Alahi, Antonio Agudo arxiv

Generative models often treat continuous data and discrete events as separate processes, creating a gap in modeling complex systems where they interact synchronously. To bridge this gap, we introduce JointDiff, a novel diffusion framework designed to unify these two processes by simultaneously generating continuous spatio-temporal data and synchronous discrete events. We demonstrate its efficacy in the sports domain by simultaneously modeling multi-agent trajectories and key possession events. This joint modeling is validated with non-controllable generation and two novel controllable generation scenarios: weak-possessor-guidance, which offers flexible semantic control over game dynamics through a simple list of intended ball possessors, and text-guidance, which enables fine-grained, language-driven generation. To enable the conditioning with these guidance signals, we introduce CrossGuid, an effective conditioning operation for multi-agent domains. We also share a new unified sports benchmark enhanced with textual descriptions for soccer and football datasets. JointDiff achieves state-of-the-art performance, demonstrating that joint modeling is crucial for building realistic and controllable generative models for interactive systems. https://guillem-cf.github.io/JointDiff/

📄 PDF Abstract BibTeX arXiv:2509.22522

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

A Versatile and Differentiable Hand-Object Interaction Representation

2024-09-25 · Théo Morales, Omid Taheri, Gerard Lacey

Synthesizing accurate hands-object interactions (HOI) is critical for applications in Computer Vision, Augmented Reality (AR), and Mixed Reality (MR). Despite recent advances, the accuracy of reconstructed or generated H…

Mixed RealityObject

Bridging the Gap Between Learning in Discrete and Continuous Environments for Vision-and-Language Navigation

2022-03-05 · CVPR 2022 1 · Yicong Hong, Zun Wang, Qi Wu, Stephen Gould

Most existing works in vision-and-language navigation (VLN) focus on either discrete or continuous environments, training agents that cannot generalize across the two. The fundamental difference between the two setups is…

Imitation LearningVision and Language Navigation

Towards Learning to Speak and Hear Through Multi-Agent Communication over a Continuous Acoustic Channel

2021-11-04 · Kevin Eloff, Okko Räsänen, Herman A. Engelbrecht, Arnu Pretorius 외

Multi-agent reinforcement learning has been used as an effective means to study emergent communication between agents, yet little focus has been given to continuous acoustic communication. This would be more akin to huma…

Language AcquisitionMulti-agent Reinforcement LearningQ-Learningreinforcement-learning+1

Physics-guided and fabrication-aware inverse design of photonic devices using diffusion models

2025-04-23 · Dongjin Seo, Soobin Um, Sangbin Lee, Jong Chul Ye 외

Designing free-form photonic devices is fundamentally challenging due to the vast number of possible geometries and the complex requirements of fabrication constraints. Traditional inverse-design approaches--whether driv…

BinarizationDenoisingglobal-optimization

Bridging Continuous and Discrete Tokens for Autoregressive Visual Generation

2025-03-20 · Yuqing Wang, Zhijie Lin, Yao Teng, Yuanzhi Zhu 외

Autoregressive visual generation models typically rely on tokenizers to compress images into tokens that can be predicted sequentially. A fundamental dilemma exists in token representation: discrete tokens enable straigh…

Quantization