paper-with-me

Papers

Exploring Vision-Language Models for Open-Vocabulary Zero-Shot Action Segmentation

2026-02-24 · Asim Unmesh, Kaki Ramesh, Mayank Patel, Rahul Jain, Karthik Ramani arxiv

Temporal Action Segmentation (TAS) requires dividing videos into action segments, yet the vast space of activities and alternative breakdowns makes collecting comprehensive datasets infeasible. Existing methods remain limited to closed vocabularies and fixed label sets. In this work, we explore the largely unexplored problem of Open-Vocabulary Zero-Shot Temporal Action Segmentation (OVTAS) by leveraging the strong zero-shot capabilities of Vision-Language Models (VLMs). We introduce a training-free pipeline that follows a segmentation-by-classification design: Frame-Action Embedding Similarity (FAES) matches video frames to candidate action labels, and Similarity-Matrix Temporal Segmentation (SMTS) enforces temporal consistency. Beyond proposing OVTAS, we present a systematic study across 14 diverse VLMs, providing the first broad analysis of their suitability for open-vocabulary action segmentation. Experiments on standard benchmarks show that OVTAS achieves strong results without task-specific supervision, underscoring the potential of VLMs for structured temporal understanding.

📄 PDF Abstract BibTeX arXiv:2602.21406

Code (0)

등록된 구현이 없습니다.

Tasks

Action Segmentation

Similar Papers 제목 키워드 기반

Exploring Open-Vocabulary Semantic Segmentation from CLIP Vision Encoder Distillation Only

2023-01-01 · ICCV 2023 1 · Jun Chen, Deyao Zhu, Guocheng Qian, Bernard Ghanem 외

Semantic segmentation is a crucial task in computer vision that involves segmenting images into semantically meaningful regions at the pixel level. However, existing approaches often rely on expensive human annotatio…

Open Vocabulary Semantic SegmentationOpen-Vocabulary Semantic SegmentationSegmentationSemantic Segmentation+2

Exploring Open-Vocabulary Semantic Segmentation without Human Labels

2023-06-01 · Jun Chen, Deyao Zhu, Guocheng Qian, Bernard Ghanem 외

Semantic segmentation is a crucial task in computer vision that involves segmenting images into semantically meaningful regions at the pixel level. However, existing approaches often rely on expensive human annotations a…

Open Vocabulary Semantic SegmentationOpen-Vocabulary Semantic SegmentationSegmentationSemantic Segmentation+2

RSVG-ZeroOV: Exploring a Training-Free Framework for Zero-Shot Open-Vocabulary Visual Grounding in Remote Sensing Images

2025-09-23 · Ke Li, Di Wang, Ting Wang, Fuyu Dong 외 arxiv

Remote sensing visual grounding (RSVG) aims to localize objects in remote sensing images based on free-form natural language expressions. Existing approaches are typically constrained to closed-set vocabularies, limiting…

Visual Grounding

Synthetic Captions for Open-Vocabulary Zero-Shot Segmentation

2025-09-15 · Tim Lebailly, Vijay Veerabadran, Satwik Kottur, Karl Ridgeway 외 arxiv

Generative vision-language models (VLMs) exhibit strong high-level image understanding but lack spatially dense alignment between vision and language modalities, as our findings indicate. Orthogonal to advancements in ge…

Representation Learning

Open-vocabulary Attribute Detection

2022-11-23 · CVPR 2023 1 · María A. Bravo, Sudhanshu Mittal, Simon Ging, Thomas Brox

Vision-language modeling has enabled open-vocabulary tasks where predictions can be queried using any text prompt in a zero-shot manner. Existing open-vocabulary tasks focus on object classes, whereas research on object …

AttributeLanguage ModelingLanguage ModellingObject+2