paper-with-me

Papers

Simplex-Based 3D Spatio-Temporal Feature Description for Action Recognition

2014-06-01 · CVPR 2014 6 · Hao Zhang, Wenjun Zhou, Christopher Reardon, Lynne E. Parker

We present a novel feature description algorithm to describe 3D local spatio-temporal features for human action recognition. Our descriptor avoids the singularity and limited discrimination power issues of traditional 3D descriptors by quantizing and describing visual features in the simplex topological vector space. Specifically, given a feature's support region containing a set of 3D visual cues, we decompose the cues' orientation into three angles, transform the decomposed angles into the simplex space, and describe them in such a space. Then, quadrant decomposition is performed to improve discrimination, and a final feature vector is composed from the resulting histograms. We develop intuitive visualization tools for analyzing feature characteristics in the simplex topological vector space. Experimental results demonstrate that our novel simplex-based orientation decomposition (SOD) descriptor substantially outperforms traditional 3D descriptors for the KTH, UCF Sport, and Hollywood-2 benchmark action datasets. In addition, the results show that our SOD descriptor is a superior individual descriptor for action recognition.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Action RecognitionTemporal Action Localization

Similar Papers 제목 키워드 기반

Spatio-Temporal Ranked-Attention Networks for Video Captioning

2020-01-17 · Anoop Cherian, Jue Wang, Chiori Hori, Tim K. Marks

Generating video descriptions automatically is a challenging task that involves a complex interplay between spatio-temporal visual features and language models. Given that videos consist of spatial (frame-level) features…

Video Captioning

Caption Anything in Video: Fine-grained Object-centric Captioning via Spatiotemporal Multimodal Prompting

2025-04-07 · Yunlong Tang, Jing Bi, Chao Huang, Susan Liang 외

We present CAT-V (Caption AnyThing in Video), a training-free framework for fine-grained object-centric video captioning that enables detailed descriptions of user-selected objects through time. CAT-V integrates three ke…

Boundary DetectionObjectSemantic SegmentationVideo Captioning

Spatiotemporal Prediction of Electric Vehicle Charging Load Based on Large Language Models

2025-06-04 · Hang Fan, Mingxuan Li, Jingshi Cui, Zuhan Zhang 외

The rapid growth of EVs and the subsequent increase in charging demand pose significant challenges for load grid scheduling and the operation of EV charging stations. Effectively harnessing the spatiotemporal correlation…

Scheduling

FineMoGen: Fine-Grained Spatio-Temporal Motion Generation and Editing

2023-12-22 · NeurIPS 2023 11 · Mingyuan Zhang, Huirong Li, Zhongang Cai, Jiawei Ren 외

Text-driven motion generation has achieved substantial progress with the emergence of diffusion models. However, existing methods still struggle to generate complex motion sequences that correspond to fine-grained descri…

Mixture-of-ExpertsMotion GenerationMotion Synthesis

Mining Interpretable Spatio-temporal Logic Properties for Spatially Distributed Systems

2021-06-16 · Sara Mohammadinejad, Jyotirmy V. Deshmukh, Laura Nenzi

The Internet-of-Things, complex sensor networks, multi-agent cyber-physical systems are all examples of spatially distributed systems that continuously evolve in time. Such systems generate huge amounts of spatio-tempora…

ClusteringEpidemiology