paper-with-me

홈 › Papers

Storyboard guided Alignment for Fine-grained Video Action Recognition

2024-10-18 · Enqi Liu, Liyuan Pan, Yan Yang, Yiran Zhong, Zhijing Wu, Xinxiao wu, Liu Liu

Fine-grained video action recognition can be conceptualized as a video-text matching problem. Previous approaches often rely on global video semantics to consolidate video embeddings, which can lead to misalignment in video-text pairs due to a lack of understanding of action semantics at an atomic granularity level. To tackle this challenge, we propose a multi-granularity framework based on two observations: (i) videos with different global semantics may share similar atomic actions or appearances, and (ii) atomic actions within a video can be momentary, slow, or even non-directly related to the global video semantics. Inspired by the concept of storyboarding, which disassembles a script into individual shots, we enhance global video semantics by generating fine-grained descriptions using a pre-trained large language model. These detailed descriptions capture common atomic actions depicted in videos. A filtering metric is proposed to select the descriptions that correspond to the atomic actions present in both the videos and the descriptions. By employing global semantics and fine-grained descriptions, we can identify key frames in videos and utilize them to aggregate embeddings, thereby making the embedding more accurate. Extensive experiments on various video action recognition datasets demonstrate superior performance of our proposed method in supervised, few-shot, and zero-shot settings.

📄 PDF Abstract BibTeX arXiv:2410.14238

Code (0)

등록된 구현이 없습니다.

Tasks

Action RecognitionLanguage ModellingLarge Language ModelTemporal Action LocalizationText Matching

Similar Papers 제목 키워드 기반

STAGE: Storyboard-Anchored Generation for Cinematic Multi-shot Narrative

2025-12-13 · Peixuan Zhang, Zijian Jia, Kaiqi Liu, Shuchen Weng 외 arxiv

While recent advancements in generative models have achieved remarkable visual fidelity in video synthesis, creating coherent multi-shot narratives remains a significant challenge. To address this, keyframe-based approac…

Video Generation

DrawVideo: Generating Long Video from Storyboard Keyframe Sketches

2026-05-22 · Chuanzhi Xu, Huiqi Liang, Bang Shi, Huiming Zhang 외 arxiv

Long video generation requires high-fidelity synthesis, coherent narrative structure, and user control over extended time spans. Existing text-to-video methods often rely on a single long prompt, limiting control over po…

Video Generation

DreamShot: Personalized Storyboard Synthesis with Video Diffusion Prior

2026-04-19 · Junjia Huang, Binbin Yang, Pengxiang Yan, Jiyang Liu 외 arxiv

Storyboard synthesis plays a crucial role in visual storytelling, aiming to generate coherent shot sequences that visually narrate cinematic events with consistent characters, scenes, and transitions. However, existing a…

Visual StorytellingStory Continuation

Dynamic Storyboard Generation in an Engine-based Virtual Environment for Video Production

2023-01-30 · Anyi Rao, Xuekun Jiang, Yuwei Guo, Linning Xu 외

Amateurs working on mini-films and short-form videos usually spend lots of time and effort on the multi-round complicated process of setting and adjusting scenes, plots, and cameras to deliver satisfying video shots. We …

ActAlign: Zero-Shot Fine-Grained Video Classification via Language-Guided Sequence Alignment

2025-06-28 · Amir Aghdam, Vincent Tao Hu

We address the task of zero-shot fine-grained video classification, where no video examples or temporal annotations are available for unseen action classes. While contrastive vision-language models such as SigLIP demonst…

Dynamic Time WarpingLarge Language ModelOpen Set Learningtext similarity+2