paper-with-me

홈 › Papers

Modeling Cross-view Interaction Consistency for Paired Egocentric Interaction Recognition

2020-03-24 · Zhongguo Li, Fan Lyu, Wei Feng, Song Wang

With the development of Augmented Reality (AR), egocentric action recognition (EAR) plays important role in accurately understanding demands from the user. However, EAR is designed to help recognize human-machine interaction in single egocentric view, thus difficult to capture interactions between two face-to-face AR users. Paired egocentric interaction recognition (PEIR) is the task to collaboratively recognize the interactions between two persons with the videos in their corresponding views. Unfortunately, existing PEIR methods always directly use linear decision function to fuse the features extracted from two corresponding egocentric videos, which ignore consistency of interaction in paired egocentric videos. The consistency of interactions in paired videos, and features extracted from them are correlated to each other. On top of that, we propose to build the relevance between two views using biliear pooling, which capture the consistency of two views in feature-level. Specifically, each neuron in the feature maps from one view connects to the neurons from another view, which guarantee the compact consistency between two views. Then all possible paired neurons are used for PEIR for the inside consistent information of them. To be efficient, we use compact bilinear pooling with Count Sketch to avoid directly computing outer product in bilinear. Experimental results on dataset PEV shows the superiority of the proposed methods on the task PEIR.

📄 PDF Abstract BibTeX arXiv:2003.10663

Code (0)

등록된 구현이 없습니다.

Tasks

Action Recognition

Similar Papers 제목 키워드 기반

ShareVerse: Multi-Agent Consistent Video Generation for Shared World Modeling

2026-03-03 · Jiayi Zhu, Jianing Zhang, Yiying Yang, Wei Cheng 외 arxiv

This paper presents ShareVerse, a video generation framework enabling multi-agent shared world modeling, addressing the gap in existing works that lack support for unified shared world construction with multi-agent inter…

Video Generation

View-Consistent 3D Scene Editing via Dual-Path Structural Correspondense and Semantic Continuity

2026-04-20 · Pufan Li, Bi'an Du, Shenghe Zheng, Junyi Yao 외 arxiv

Text-driven 3D scene editing has recently attracted increasing attention. Most existing methods follow a render-edit-optimize pipeline, where multi-view images are rendered from a 3D scene, edited with 2D image editors, …

3D scene Editing

Identity-Consistent Video Generation under Large Facial-Angle Variations

2026-03-22 · Bin Hu, Zipeng Qi, Guoxi Huang, Zunnan Xu 외 arxiv

Single-view reference-to-video methods often struggle to preserve identity consistency under large facial-angle variations. This limitation naturally motivates the incorporation of multi-view facial references. However, …

Video Generation

Learning Reactive Human Motion Generation from Paired Interaction Data Using Transformer-Based Models

2026-04-24 · Masato Soga, Ryuki Takebayashi arxiv

Recent advances in deep learning have enabled the generation of videos from textual descriptions as well as the prediction of future sequences from input videos. Similarly, in human motion modeling, motions can be genera…

Inclusive Interactive Collisions for Multi-View Consistent Compositional 3D Generation

2026-06-23 · Chang Liu, Mingwen Shao, Xiang Lv, Xinyuan Chen 외 arxiv

Recent breakthroughs in 3D generation have advanced notably with the development of text-to-image diffusion model. However, existing methods remain two practical challenges: (1) They primarily generate single 3D object, …

Scene Generation3D Generation