paper-with-me

Papers

Identity-Aware Human-Object Interaction Motion Captioning

2026-08-21 · Yiming Wang, Yonghao Dang, Huilai Li, Jiawei Tu, Jianqin Yin arxiv

Existing human-object interaction (HOI) motion captioning methods typically describe what happens while referring to the subject using generic terms such as "a person" or "someone", without grounding the caption in subject identity. To address this limitation, we introduce Identity-Aware Human-Object Interaction Motion Captioning task. This task requires each generated caption to specify both the subject identity and the corresponding HOI motion. For example, the model generates "Sub_ID lifts the chair" rather than "A person lifts the chair". For this task, we design identity-aware HOI motion captions based on the BEHAVE and InterCap datasets. We further propose ID-HOINet, which learns from multi-view videos while supporting single-view identity-aware HOI motion caption generation. ID-HOINet contains two core components: Multi-View Identity-Motion Learning Module (MVIML) and Two-Stage Caption Rewriting Strategy (TSCR). MVIML learns from multi-view videos by modeling dependencies across temporal stages and camera viewpoints, capturing identity and interaction motion features. At inference, the TSCR first retrieves the subject identity and generates identity-agnostic HOI motion captions. TSCR then rewrites these captions with the predicted identity to produce the final identity-aware HOI motion captions. Experiments demonstrate that ID-HOINet achieves state-of-the-art performance. Code will be released upon acceptance.

📄 PDF Abstract BibTeX arXiv:2608.20690

Code (0)

등록된 구현이 없습니다.

Tasks

Motion Captioning

Similar Papers 제목 키워드 기반

Vera: Identity-Faithful Human Subject-to-Video Generation

2026-07-22 · Yulong Xu, Xinyue Liu, Shujuan Li, huafeng shi 외 arxiv

Subject-to-video (S2V) generation has made substantial progress in preserving reference subjects across diverse categories, yet generic subject consistency remains insufficient for human-centric generation. A video may a…

Video Generation

PA-HOI: A Physics-Aware Human and Object Interaction Dataset

2025-08-08 · Ruiyan Wang, Lin Zuo, Zonghao Lin, Qiang Wang 외 arxiv

The Human-Object Interaction (HOI) task explores the dynamic interactions between humans and objects in physical environments, providing essential biomechanical and cognitive-behavioral foundations for fields such as rob…

Object-Aware 4D Human Motion Generation

2025-10-31 · Shurui Gui, Deep Anil Patel, Xiner Li, Martin Renqiang Min arxiv

Recent advances in video diffusion models have enabled the generation of high-quality videos. However, these videos still suffer from unrealistic deformations, semantic violations, and physical inconsistencies that are l…

Context-aware Human Motion Prediction

2019-04-06 · CVPR 2020 6 · Enric Corona, Albert Pumarola, Guillem Alenyà, Francesc Moreno-Noguer

The problem of predicting human motion given a sequence of past observations is at the core of many applications in robotics and computer vision. Current state-of-the-art formulate this problem as a sequence-to-sequence …

Graph AttentionHuman motion predictionmotion predictionPrediction

LSF-Animation: Label-Free Speech-Driven Facial Animation via Implicit Feature Representation

2025-10-23 · Xin Lu, Chuanqing Zhuang, Chenxi Jin, Zhengda Lu 외 arxiv

Speech-driven 3D facial animation has attracted increasing interest since its potential to generate expressive and temporally synchronized digital humans. While recent works have begun to explore emotion-aware animation,…