paper-with-me

Papers

Knowledge Distillation for Multimodal Egocentric Action Recognition Robust to Missing Modalities

2025-04-11 · Maria Santos-Villafranca, Dustin Carrión-Ojeda, Alejandro Perez-Yus, Jesus Bermudez-Cameo, Jose J. Guerrero, Simone Schaub-Meyer

Action recognition is an essential task in egocentric vision due to its wide range of applications across many fields. While deep learning methods have been proposed to address this task, most rely on a single modality, typically video. However, including additional modalities may improve the robustness of the approaches to common issues in egocentric videos, such as blurriness and occlusions. Recent efforts in multimodal egocentric action recognition often assume the availability of all modalities, leading to failures or performance drops when any modality is missing. To address this, we introduce an efficient multimodal knowledge distillation approach for egocentric action recognition that is robust to missing modalities (KARMMA) while still benefiting when multiple modalities are available. Our method focuses on resource-efficient development by leveraging pre-trained models as unimodal feature extractors in our teacher model, which distills knowledge into a much smaller and faster student model. Experiments on the Epic-Kitchens and Something-Something datasets demonstrate that our student model effectively handles missing modalities while reducing its accuracy drop in this scenario.

📄 PDF Abstract BibTeX arXiv:2504.08578

Code (0)

등록된 구현이 없습니다.

Tasks

Action RecognitionKnowledge Distillation

Methods 이 논문이 사용한 방법론

Knowledge Distillation A very simple way to improve the performance of almost any machine learning algorithm is to train many different models on the same data and then to average their predictions.…

Similar Papers 제목 키워드 기반

Multimodal Distillation for Egocentric Action Recognition

2023-07-14 · ICCV 2023 1 · Gorjan Radevski, Dusan Grujicic, Marie-Francine Moens, Matthew Blaschko 외

The focal point of egocentric video understanding is modelling hand-object interactions. Standard models, e.g. CNNs or Vision Transformers, which receive RGB frames as input perform well. However, their performance impro…

Action RecognitionKnowledge DistillationOptical Flow EstimationVideo Understanding

Multimodal Cross-Domain Few-Shot Learning for Egocentric Action Recognition

2024-05-30 · Masashi Hatano, Ryo Hachiuma, Ryo Fujii, Hideo Saito

We address a novel cross-domain few-shot learning task (CD-FSL) with multimodal input and unlabeled target data for egocentric action recognition. This paper simultaneously tackles two critical challenges associated with…

Action RecognitionCross-Domain Few-Shotcross-domain few-shot learningFew-Shot Learning

Students taught by multimodal teachers are superior action recognizers

2022-10-09 · Gorjan Radevski, Dusan Grujicic, Matthew Blaschko, Marie-Francine Moens 외

The focal point of egocentric video understanding is modelling hand-object interactions. Standard models -- CNNs, Vision Transformers, etc. -- which receive RGB frames as input perform well, however, their performance im…

Action RecognitionKnowledge DistillationObjectOptical Flow Estimation+1

UNIEGO: Proxies as Mediators for Unified Egocentric Video Representation Learning

2026-06-18 · Wenhao Chi, Arkaprava Sinha, Dominick Reilly, Hieu Le 외 arxiv

Egocentric video understanding is inherently limited by the narrow perspective of wearable cameras: a single viewpoint, a single modality, a single model cannot capture the full richness of human action. We argue that a …

Representation LearningAction SegmentationAction RecognitionVideo Retrieval

Bridging Modalities and Transferring Knowledge: Enhanced Multimodal Understanding and Recognition

2025-12-23 · Gorjan Radevski arxiv

This manuscript explores multimodal alignment, translation, fusion, and transference to enhance machine understanding of complex inputs. We organize the work into five chapters, each addressing unique challenges in multi…

Knowledge DistillationAction RecognitionObject DetectionKnowledge Graphs