paper-with-me

Papers

Learning multimodal representations for sample-efficient recognition of human actions

2019-03-06 · Miguel Vasco, Francisco S. Melo, David Martins de Matos, Ana Paiva, Tetsunari Inamura

Humans interact in rich and diverse ways with the environment. However, the representation of such behavior by artificial agents is often limited. In this work we present \textit{motion concepts}, a novel multimodal representation of human actions in a household environment. A motion concept encompasses a probabilistic description of the kinematics of the action along with its contextual background, namely the location and the objects held during the performance. Furthermore, we present Online Motion Concept Learning (OMCL), a new algorithm which learns novel motion concepts from action demonstrations and recognizes previously learned motion concepts. The algorithm is evaluated on a virtual-reality household environment with the presence of a human avatar. OMCL outperforms standard motion recognition algorithms on an one-shot recognition task, attesting to its potential for sample-efficient recognition of human actions.

📄 PDF Abstract BibTeX arXiv:1903.02511

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Multimodal Language Analysis with Recurrent Multistage Fusion

2018-08-12 · EMNLP 2018 10 · Paul Pu Liang, Ziyin Liu, Amir Zadeh, Louis-Philippe Morency

Computational modeling of human multimodal language is an emerging research area in natural language processing spanning the language, visual and acoustic modalities. Comprehending multimodal language requires modeling n…

Emotion RecognitionMultimodal Sentiment AnalysisSentiment Analysis

WavFusion: Towards wav2vec 2.0 Multimodal Speech Emotion Recognition

2024-12-07 · Feng Li, Jiusong Luo, Wanjun Xia

Speech emotion recognition (SER) remains a challenging yet crucial task due to the inherent complexity and diversity of human emotions. To address this problem, researchers attempt to fuse information from other modaliti…

DiversityEmotion RecognitionRepresentation LearningSpeech Emotion Recognition

Spoken in Jest, Detected in Earnest: A Systematic Review of Sarcasm Recognition -- Multimodal Fusion, Challenges, and Future Prospects

2025-09-04 · Xiyuan Gao, Shekhar Nayak, Matt Coler arxiv

Sarcasm, a common feature of human communication, poses challenges in interpersonal interactions and human-machine interactions. Linguistic research has highlighted the importance of prosodic cues, such as variations in …

Sarcasm Detection

MIntRec2.0: A Large-scale Benchmark Dataset for Multimodal Intent Recognition and Out-of-scope Detection in Conversations

2024-03-16 · Hanlei Zhang, Xin Wang, Hua Xu, Qianrui Zhou 외

Multimodal intent recognition poses significant challenges, requiring the incorporation of non-verbal modalities from real-world contexts to enhance the comprehension of human intentions. Existing benchmark datasets are …

Intent RecognitionMultimodal Intent Recognition

Geometric Visual Fusion Graph Neural Networks for Multi-Person Human-Object Interaction Recognition in Videos

2025-06-03 · Tanqiu Qiao, Ruochen Li, Frederick W. B. Li, Yoshiki Kubotani 외

Human-Object Interaction (HOI) recognition in videos requires understanding both visual patterns and geometric relationships as they evolve over time. Visual and geometric features offer complementary strengths. Visual f…

Graph LearningGraph Neural NetworkHuman-Object Interaction Detection