paper-with-me

Papers

COLLAGE: Collaborative Human-Agent Interaction Generation using Hierarchical Latent Diffusion and Language Models

2024-09-30 · Divyanshu Daiya, Damon Conover, Aniket Bera

We propose a novel framework COLLAGE for generating collaborative agent-object-agent interactions by leveraging large language models (LLMs) and hierarchical motion-specific vector-quantized variational autoencoders (VQ-VAEs). Our model addresses the lack of rich datasets in this domain by incorporating the knowledge and reasoning abilities of LLMs to guide a generative diffusion model. The hierarchical VQ-VAE architecture captures different motion-specific characteristics at multiple levels of abstraction, avoiding redundant concepts and enabling efficient multi-resolution representation. We introduce a diffusion model that operates in the latent space and incorporates LLM-generated motion planning cues to guide the denoising process, resulting in prompt-specific motion generation with greater control and diversity. Experimental results on the CORE-4D, and InterHuman datasets demonstrate the effectiveness of our approach in generating realistic and diverse collaborative human-object-human interactions, outperforming state-of-the-art methods. Our work opens up new possibilities for modeling complex interactions in various domains, such as robotics, graphics and computer vision.

📄 PDF Abstract BibTeX arXiv:2409.20502

Code (0)

등록된 구현이 없습니다.

Tasks

DenoisingDiversityMotion GenerationMotion Planning

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…
VQ-VAE VQ-VAE is a type of variational autoencoder that uses vector quantisation to obtain a discrete latent representation. It differs from…

Similar Papers 제목 키워드 기반

Neural Collage Transfer: Artistic Reconstruction via Material Manipulation

2023-11-03 · ICCV 2023 1 · Ganghun Lee, Minji Kim, Yunsu Lee, Minsu Lee 외

Collage is a creative art form that uses diverse material scraps as a base unit to compose a single image. Although pixel-wise generation techniques can reproduce a target image in collage style, it is not a suitable met…

Aesthetic Photo Collage with Deep Reinforcement Learning

2021-10-19 · Mingrui Zhang, Mading Li, Li Chen, Jiahao Yu

Photo collage aims to automatically arrange multiple photos on a given canvas with high aesthetic quality. Existing methods are based mainly on handcrafted feature optimization, which cannot adequately capture high-level…

Deep Reinforcement Learningreinforcement-learningReinforcement LearningReinforcement Learning (RL)

Semantic Prompting: Agentic Incremental Narrative Refinement through Spatial Semantic Interaction

2026-04-21 · Xuxin Tang, Ibrahim Tahmid, Eric Krokos, Kirsten Whitley 외 arxiv

Interactive spatial layouts empower users to synthesize information and organize findings for sensemaking. While Large Language Models (LLMs) can automate narrative generation from spatial layouts, current collage-based …

AgentCF: Collaborative Learning with Autonomous Language Agents for Recommender Systems

2023-10-13 · Junjie Zhang, Yupeng Hou, Ruobing Xie, Wenqi Sun 외

Recently, there has been an emergence of employing LLM-powered agents as believable human proxies, based on their remarkable decision-making capability. However, existing studies mainly focus on simulating human dialogue…

Collaborative FilteringDecision MakingLanguage ModelingLanguage Modelling+1

Self-Reasoning Agentic Framework for Narrative Product Grid-Collage Generation

2026-04-18 · Minyan Luo, Yuxin Zhang, Yifei Li, Xincan Wang 외 arxiv

Narrative-driven product photography has become a prevalent paradigm in modern marketing, as coherent visual storytelling helps convey product value and establishes emotional engagement with consumers. However, existing …

Visual StorytellingImage Generation