paper-with-me

Papers

GHOST: Grounded Human Motion Generation with Open Vocabulary Scene-and-Text Contexts

2024-04-08 · Zoltán Á. Milacski, Koichiro Niinuma, Ryosuke Kawamura, Fernando de la Torre, László A. Jeni

The connection between our 3D surroundings and the descriptive language that characterizes them would be well-suited for localizing and generating human motion in context but for one problem. The complexity introduced by multiple modalities makes capturing this connection challenging with a fixed set of descriptors. Specifically, closed vocabulary scene encoders, which require learning text-scene associations from scratch, have been favored in the literature, often resulting in inaccurate motion grounding. In this paper, we propose a method that integrates an open vocabulary scene encoder into the architecture, establishing a robust connection between text and scene. Our two-step approach starts with pretraining the scene encoder through knowledge distillation from an existing open vocabulary semantic image segmentation model, ensuring a shared text-scene feature space. Subsequently, the scene encoder is fine-tuned for conditional motion generation, incorporating two novel regularization losses that regress the category and size of the goal object. Our methodology achieves up to a 30% reduction in the goal object distance metric compared to the prior state-of-the-art baseline model on the HUMANISE dataset. This improvement is demonstrated through evaluations conducted using three implementations of our framework and a perceptual study. Additionally, our method is designed to seamlessly accommodate future 2D segmentation methods that provide per-pixel text-aligned features for distillation.

📄 PDF Abstract BibTeX arXiv:2405.18438

Code (0)

등록된 구현이 없습니다.

Tasks

DescriptiveImage SegmentationKnowledge DistillationMotion GenerationSemantic Segmentation

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically
Knowledge Distillation A very simple way to improve the performance of almost any machine learning algorithm is to train many different models on the same data and then to average their predictions.…

Similar Papers 제목 키워드 기반

From Language to Locomotion: Retargeting-free Humanoid Control via Motion Latent Guidance

2025-10-16 · Zhe Li, Cheng Chi, Yangyang Wei, Boan Zhu 외 arxiv

Natural language offers a natural interface for humanoid robots, but existing language-guided humanoid locomotion pipelines remain cumbersome and untrustworthy. They typically decode human motion, retarget it to robot mo…

Making Avatars Interact: Towards Text-Driven Human-Object Interaction for Controllable Talking Avatars

2026-02-02 · Youliang Zhang, Zhengguang Zhou, Zhentao Yu, Ziyao Huang 외 arxiv

Generating talking avatars is a fundamental task in video generation. Although existing methods can generate full-body talking avatars with simple human motion, extending this task to grounded human-object interaction (G…

Video Generation

The AI Ghostwriter Effect: When Users Do Not Perceive Ownership of AI-Generated Text But Self-Declare as Authors

2023-03-06 · Fiona Draxler, Anna Werner, Florian Lehmann, Matthias Hoppe 외

Human-AI interaction in text production increases complexity in authorship. In two empirical studies (n1 = 30 & n2 = 96), we investigate authorship and ownership in human-AI collaboration for personalized language genera…

AttributeText Generation

Grounded Gesture Generation: Language, Motion, and Space

2025-07-06 · Anna Deichler, Jim O'Regan, Teo Guichoux, David Johansson 외 arxiv

Human motion generation has advanced rapidly in recent years, yet the critical problem of creating spatially grounded, context-aware gestures has been largely overlooked. Existing models typically specialize either in de…

Synthetic Data GenerationGesture Generation

Chat-Ghosting: A Comparative Study of Methods for Auto-Completion in Dialog Systems

2025-07-08 · Sandeep Mishra, Anubhab Mandal, Bishal Santra, Tushar Abhishek 외

Ghosting, the ability to predict a user's intended text input for inline query auto-completion, is an invaluable feature for modern search engines and chat interfaces, greatly enhancing user experience. By suggesting com…

Deep Learning