GZSL Video Classification
6개 벤치마크 · 논문 7편 · 이 태스크의 논문 보기 →
Benchmarks
ActivityNet-GZSL(main)
UCF-GZSL(main)
VGGSound-GZSL(main)
ActivityNet-GZSL (cls)
UCF-GZSL (cls)
VGGSound-GZSL (cls)
Most implemented
Temporal and cross-modal attention for audio-visual zero-shot learning
Boosting Audio-visual Zero-shot Learning with Large Language Models
Papers
Boosting Audio-visual Zero-shot Learning with Large Language Models
Audio-visual zero-shot learning aims to recognize unseen classes based on paired audio-visual sequences. Recent methods mainly focus on learning multi-modal features aligned with class names to enhance the generalization…
audio-visual learningDescriptiveGZSL Video ClassificationZero-Shot LearningHyperbolic Audio-visual Zero-shot Learning
Audio-visual zero-shot learning aims to classify samples consisting of a pair of corresponding audio and video sequences from classes that are not present during training. An analysis of the audio-visual data reveals a l…
GZSL Video ClassificationZero-Shot LearningTemporal and cross-modal attention for audio-visual zero-shot learning
Audio-visual generalised zero-shot learning for video classification requires understanding the relations between the audio and visual information in order to be able to recognise samples from novel, previously unseen cl…
GZSL Video ClassificationVideo ClassificationZero-Shot LearningAttribute Prototype Network for Any-Shot Learning
Any-shot image classification allows to recognize novel classes with only a few or even zero samples. For the task of zero-shot learning, visual attributes have been shown to play an important role, while in the few-shot…
AttributeFew-Shot Image ClassificationGZSL Video Classificationimage-classification+3Audio-visual Generalised Zero-shot Learning with Cross-modal Attention and Language
Learning to classify video data from classes not included in the training data, i.e. video-based zero-shot learning, is challenging. We conjecture that the natural alignment between the audio and visual modalities in vid…
GZSL Video ClassificationZero-Shot LearningZSL Video ClassificationAVGZSLNet: Audio-Visual Generalized Zero-Shot Learning by Reconstructing Label Features from Multi-Modal Embeddings
In this paper, we propose a novel approach for generalized zero-shot learning in a multi-modal setting, where we have novel classes of audio/video during testing that are not seen during training. We use the semantic rel…
DecoderGeneralized Zero-Shot LearningGZSL Video ClassificationRetrieval+3