paper-with-me

홈 › Papers

OST: Refining Text Knowledge with Optimal Spatio-Temporal Descriptor for General Video Recognition

2023-11-30 · CVPR 2024 1 · Tongjia Chen, Hongshan Yu, Zhengeng Yang, Zechuan Li, Wei Sun, Chen Chen

Due to the resource-intensive nature of training vision-language models on expansive video data, a majority of studies have centered on adapting pre-trained image-language models to the video domain. Dominant pipelines propose to tackle the visual discrepancies with additional temporal learners while overlooking the substantial discrepancy for web-scaled descriptive narratives and concise action category names, leading to less distinct semantic space and potential performance limitations. In this work, we prioritize the refinement of text knowledge to facilitate generalizable video recognition. To address the limitations of the less distinct semantic space of category names, we prompt a large language model (LLM) to augment action class names into Spatio-Temporal Descriptors thus bridging the textual discrepancy and serving as a knowledge base for general recognition. Moreover, to assign the best descriptors with different video instances, we propose Optimal Descriptor Solver, forming the video recognition problem as solving the optimal matching flow across frame-level representations and descriptors. Comprehensive evaluations in zero-shot, few-shot, and fully supervised video recognition highlight the effectiveness of our approach. Our best model achieves a state-of-the-art zero-shot accuracy of 75.1% on Kinetics-600.

📄 PDF Abstract BibTeX arXiv:2312.00096

Code (1)

tomchen-ctj/OST 공식 구현 pytorch

Tasks

DescriptiveLanguage ModellingLarge Language ModelVideo RecognitionZero-Shot Action RecognitionZero-Shot Action Recognition on HMDB51Zero-Shot Action Recognition on UCF101

Methods 이 논문이 사용한 방법론

BASE 설명 없음

Similar Papers 제목 키워드 기반

Available Transfer Capability Calculation for Wind-Integrated Power Systems Considering Wind Speed Spatiotemporal Correlation and Primal-Dual Interior Point Method

2024-09-21 · Xia-Liang Huangpu

This paper explores the intricate effects of wind power integration on the Available Transfer Capability (ATC) of power systems, emphasizing the significance of spatiotemporal correlations in wind speed. We present an in…

RS-SSM: Refining Forgotten Specifics in State Space Model for Video Semantic Segmentation

2026-03-25 · Kai Zhu, Zhenyu Cui, Zehua Zang, Jiahuan Zhou arxiv

Recently, state space models have demonstrated efficient video segmentation through linear-complexity state space compression. However, Video Semantic Segmentation (VSS) requires pixel-level spatiotemporal modeling capab…

Video Semantic SegmentationComputational EfficiencyVideo Segmentation

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP

2024-12-13 · Yating Yu, Congqi Cao, Yueran Zhang, Qinyi Lv 외

Zero-shot action recognition (ZSAR) requires collaborative multi-modal spatiotemporal understanding. However, finetuning CLIP directly for ZSAR yields suboptimal performance, given its inherent constraints in capturing e…

Action RecognitionText AugmentationZero-Shot Action Recognition

SSTKG: Simple Spatio-Temporal Knowledge Graph for Intepretable and Versatile Dynamic Information Embedding

2024-02-19 · Ruiyi Yang, Flora D. Salim, Hao Xue

Knowledge graphs (KGs) have been increasingly employed for link prediction and recommendation using real-world datasets. However, the majority of current methods rely on static data, neglecting the dynamic nature and the…

Knowledge GraphsLink PredictionPrediction

Adaptive Uncertainty-Guided Surrogates for Efficient phase field Modeling of Dendritic Solidification

2026-02-17 · Eider Garate-Perez, Kerman López de Calle-Etxabe, Oihana Garcia, Borja Calvo 외 arxiv

The high computational cost of phase field simulations remains a major limitation for predicting dendritic solidification in metals, particularly in additive manufacturing, where microstructural control is critical. This…