paper-with-me

Papers

Enhancing Spatio-Temporal Zero-shot Action Recognition with Language-driven Description Attributes

2025-10-31 · Yehna Kim, Young-Eun Kim, Seong-Whan Lee arxiv

Vision-Language Models (VLMs) have demonstrated impressive capabilities in zero-shot action recognition by learning to associate video embeddings with class embeddings. However, a significant challenge arises when relying solely on action classes to provide semantic context, particularly due to the presence of multi-semantic words, which can introduce ambiguity in understanding the intended concepts of actions. To address this issue, we propose an innovative approach that harnesses web-crawled descriptions, leveraging a large-language model to extract relevant keywords. This method reduces the need for human annotators and eliminates the laborious manual process of attribute data creation. Additionally, we introduce a spatio-temporal interaction module designed to focus on objects and action units, facilitating alignment between description attributes and video content. In our zero-shot experiments, our model achieves impressive results, attaining accuracies of 81.0%, 53.1%, and 68.9% on UCF-101, HMDB-51, and Kinetics-600, respectively, underscoring the model's adaptability and effectiveness across various downstream tasks.

📄 PDF Abstract BibTeX arXiv:2510.27255

Code (0)

등록된 구현이 없습니다.

Tasks

Zero-Shot Action Recognition

Similar Papers 제목 키워드 기반

FlowZero: Zero-Shot Text-to-Video Synthesis with LLM-Driven Dynamic Scene Syntax

2023-11-27 · Yu Lu, Linchao Zhu, Hehe Fan, Yi Yang

Text-to-video (T2V) generation is a rapidly growing research area that aims to translate the scenes, objects, and actions within complex video text into a sequence of coherent visual frames. We present FlowZero, a novel …

Video Generation

Zero-Shot Skeleton-Based Action Anticipation

2026-08-14 · Hongsong Wang, Pengbo Yan, Yang Zhang, Qiuxia Lai arxiv

Action anticipation (AA) aims to recognize ongoing human or humanoids actions from partial observations, enabling robots to predict intentions before the actions are completed. Although skeleton-based AA offers efficienc…

Zero-shot GeneralizationAction Anticipation

Interaction-Aware Prompting for Zero-Shot Spatio-Temporal Action Detection

2023-04-10 · Wei-Jhe Huang, Jheng-Hsien Yeh, Min-Hung Chen, Gueter Josmy Faure 외

The goal of spatial-temporal action detection is to determine the time and place where each person's action occurs in a video and classify the corresponding action category. Most of the existing methods adopt fully-super…

Action DetectionLanguage ModelingLanguage ModellingZero-Shot Learning

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP

2024-12-13 · Yating Yu, Congqi Cao, Yueran Zhang, Qinyi Lv 외

Zero-shot action recognition (ZSAR) requires collaborative multi-modal spatiotemporal understanding. However, finetuning CLIP directly for ZSAR yields suboptimal performance, given its inherent constraints in capturing e…

Action RecognitionText AugmentationZero-Shot Action Recognition

Skeleton-based Zero-Shot Spatio-Temporal Action Localization via Weakly-Supervised Pretraining

2026-08-26 · Koshiro Nagano, Fumiaki Sato, Ryo Hachiuma, Kazuki Tsutsukawa 외 arxiv

We propose a novel pretraining strategy for skeleton-based zero-shot spatio-temporal action localization to estimate unseen actions for person instances while overcoming high annotation costs for training via new target …

Spatio-Temporal Action LocalizationContrastive Learning