paper-with-me

홈 › Papers

SL-DML: Signal Level Deep Metric Learning for Multimodal One-Shot Action Recognition

2020-04-23 · Raphael Memmesheimer, Nick Theisen, Dietrich Paulus

Recognizing an activity with a single reference sample using metric learning approaches is a promising research field. The majority of few-shot methods focus on object recognition or face-identification. We propose a metric learning approach to reduce the action recognition problem to a nearest neighbor search in embedding space. We encode signals into images and extract features using a deep residual CNN. Using triplet loss, we learn a feature embedding. The resulting encoder transforms features into an embedding space in which closer distances encode similar actions while higher distances encode different actions. Our approach is based on a signal level formulation and remains flexible across a variety of modalities. It further outperforms the baseline on the large scale NTU RGB+D 120 dataset for the One-Shot action recognition protocol by 5.6%. With just 60% of the training data, our approach still outperforms the baseline approach by 3.7%. With 40% of the training data, our approach performs comparably well to the second follow up. Further, we show that our approach generalizes well in experiments on the UTD-MHAD dataset for inertial, skeleton and fused data and the Simitate dataset for motion capturing data. Furthermore, our inter-joint and inter-sensor experiments suggest good capabilities on previously unseen setups.

📄 PDF Abstract BibTeX arXiv:2004.11085

Code (1)

raphaelmemmesheimer/sl-dml 공식 구현 pytorch

Tasks

Action RecognitionFace IdentificationMetric LearningObject RecognitionOne-Shot 3D Action RecognitionTriplet

Similar Papers 제목 키워드 기반

Towards Efficient Multimodal and Multilingual Opinion Extraction for STI: A QLoRA-Based Fine-Tuning Approach

2026-08-14 · Sheng Hong, Xuanqi Wang, Jiacheng Wang, Yuwei Wang arxiv

Recent advances in large language models (LLMs) have reshaped semantic analysis. Opinion Extraction (OE) for Science and Technology Intelligence (STI) requires concise core opinions from large information streams. Off-th…

Multimodal Prototype-Enhanced Network for Few-Shot Action Recognition

2022-12-09 · Xinzhe Ni, Yong liu, Hao Wen, Yatai Ji 외

Current methods for few-shot action recognition mainly fall into the metric learning framework following ProtoNet, which demonstrates the importance of prototypes. Although they achieve relatively good performance, the e…

Action RecognitionFew-Shot action recognitionFew Shot Action RecognitionMetric Learning

MMRQA: Signal-Enhanced Multimodal Large Language Models for MRI Quality Assessment

2025-09-29 · Fankai Jia, Daisong Gan, Zhe Zhang, Zhaochi Wen 외 arxiv

Magnetic resonance imaging (MRI) quality assessment is crucial for clinical decision-making, yet remains challenging due to data scarcity and protocol variability. Traditional approaches face fundamental trade-offs: sign…

Zero-shot Generalization

OmniDirector: General Multi-Shot Camera Cloning without Cross-Paired Data

2026-06-11 · Jiwen Liu, Shujuan Li, Zhixue Fang, Xiaohan Li 외 arxiv

Cloning camera motion from reference videos is an important task in video generation, as videos provide intuitive and precise control. Existing methods either directly use parametric representations that fail to handle m…

Video Generation

Knowledge is Power: Advancing Few-shot Action Recognition with Multimodal Semantics from MLLMs

2026-03-27 · Jiazheng Xing, Chao Xu, Hangjie Yuan, Mengmeng Wang 외 arxiv

Multimodal Large Language Models (MLLMs) have propelled the field of few-shot action recognition (FSAR). However, preliminary explorations in this area primarily focus on generating captions to form a suboptimal feature-…

Action RecognitionMetric Learning