paper-with-me

Papers

HiSync: Spatio-Temporally Aligning Hand Motion from Wearable IMU and On-Robot Camera for Command Source Identification in Long-Range HRI

2026-03-12 · Chengwen Zhang, Chun Yu, Borong Zhuang, Haopeng Jin, Qingyang Wan, Zhuojun Li, Zhe He, Zhoutong Ye, Yu Mei, Chang Liu, Weinan Shi, Yuanchun Shi arxiv

Long-range Human-Robot Interaction (HRI) remains underexplored. Within it, Command Source Identification (CSI) - determining who issued a command - is especially challenging due to multi-user and distance-induced sensor ambiguity. We introduce HiSync, an optical-inertial fusion framework that treats hand motion as binding cues by aligning robot-mounted camera optical flow with hand-worn IMU signals. We first elicit a user-defined (N=12) gesture set and collect a multimodal command gesture dataset (N=38) in long-range multi-user HRI scenarios. Next, HiSync extracts frequency-domain hand motion features from both camera and IMU data, and a learned CSINet denoises IMU readings, temporally aligns modalities, and performs distance-aware multi-window fusion to compute cross-modal similarity of subtle, natural gestures, enabling robust CSI. In three-person scenes up to 34m, HiSync achieves 92.32% CSI accuracy, outperforming the prior SOTA by 48.44%. HiSync is also validated on real-robot deployment. By making CSI reliable and natural, HiSync provides a practical primitive and design guidance for public-space HRI. https://github.com/OctopusWen/HiSync

📄 PDF Abstract BibTeX arXiv:2603.11809

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Vectorized Video Representation with Easy Editing via Hierarchical Spatio-Temporally Consistent Proxy Embedding

2025-10-14 · Ye Chen, Liming Tan, Yupeng Zhu, Yuanbin Wang 외 arxiv

Current video representations heavily rely on unstable and over-grained priors for motion and appearance modelling, \emph{i.e.}, pixel-level matching and tracking. A tracking error of just a few pixels would lead to the …

Video Reconstruction

Estimating Ego-Body Pose from Doubly Sparse Egocentric Video Data

2024-11-05 · Seunggeun Chi, Pin-Hao Huang, Enna Sachdeva, Hengbo Ma 외

We study the problem of estimating the body movements of a camera wearer from egocentric videos. Current methods for ego-body pose estimation rely on temporally dense sensor data, such as IMU measurements from spatially …

ImputationPose Estimation

A structured latent space for human body motion generation

2021-06-07 · Mathieu Marsot, Stefanie Wuhrer, Jean-Sebastien Franco, Stephane Durocher

We propose a framework to learn a structured latent space to represent 4D human body motion, where each latent vector encodes a full motion of the whole 3D human shape. On one hand several data-driven skeletal animation …

3D geometryHuman motion predictionMotion Generationmotion prediction

A Tool for Spatio-Temporal Analysis of Social Anxiety with Twitter Data

2019-01-23 · Joohong Lee, Dongyoung Son, Yong Suk Choi

In this paper, we present a tool for analyzing spatio-temporal distribution of social anxiety. Twitter, one of the most popular social network services, has been chosen as data source for analysis of social anxiety. Twee…

HMD-NeMo: Online 3D Avatar Motion Generation From Sparse Observations

2023-08-22 · ICCV 2023 1 · Sadegh Aliakbarian, Fatemeh Saleh, David Collier, Pashmina Cameron 외

Generating both plausible and accurate full body avatar motion is the key to the quality of immersive experiences in mixed reality scenarios. Head-Mounted Devices (HMDs) typically only provide a few input signals, such a…

Mixed RealityMotion Generation