paper-with-me

Papers

Feature sampling and partitioning for visual vocabulary generation on large action classification datasets

2014-05-29 · Michael Sapienza, Fabio Cuzzolin, Philip H. S. Torr

The recent trend in action recognition is towards larger datasets, an increasing number of action classes and larger visual vocabularies. State-of-the-art human action classification in challenging video data is currently based on a bag-of-visual-words pipeline in which space-time features are aggregated globally to form a histogram. The strategies chosen to sample features and construct a visual vocabulary are critical to performance, in fact often dominating performance. In this work we provide a critical evaluation of various approaches to building a vocabulary and show that good practises do have a significant impact. By subsampling and partitioning features strategically, we are able to achieve state-of-the-art results on 5 major action recognition datasets using relatively small visual vocabularies.

📄 PDF Abstract BibTeX arXiv:1405.7545

Code (0)

등록된 구현이 없습니다.

Tasks

Action ClassificationAction RecognitionGeneral ClassificationTemporal Action Localization

Similar Papers 제목 키워드 기반

SSG: Logit-Balanced Vocabulary Partitioning for LLM Watermarking

2026-04-24 · Chenxi Gu, Xiaoning Du, John Grundy arxiv

Watermarking has emerged as a promising technique for tracing the authorship of content generated by large language models (LLMs). Among existing approaches, the KGW scheme is particularly attractive due to its versatili…

Mathematical ReasoningCode Generation

VLAD-VSA: Cross-Domain Face Presentation Attack Detection with Vocabulary Separation and Adaptation

2022-02-21 · Jiong Wang, Zhou Zhao, Weike Jin, Xinyu Duan 외

For face presentation attack detection (PAD), most of the spoofing cues are subtle, local image patterns (e.g., local image distortion, 3D mask edge and cut photo edges). The representations of existing PAD works with si…

DiversityFace Presentation Attack Detection

V2Flow: Unifying Visual Tokenization and Large Language Model Vocabularies for Autoregressive Image Generation

2025-03-10 · Guiwei Zhang, Tianyu Zhang, Mohan Zhou, Yalong Bai 외

We propose V2Flow, a novel tokenizer that produces discrete visual tokens capable of high-fidelity reconstruction, while ensuring structural and latent distribution alignment with the vocabulary space of large language m…

DecoderImage GenerationLanguage ModelingLanguage Modelling+1

Prior-SG: Task and Prior Driven Region Segmentation for Scene Graphs in Arbitrarily-Structured Environments

2026-08-06 · Giorgio Tonetti, Laurent Kneip, Abel Gawel, Marco Hutter arxiv

Hierarchical 3D scene graphs are a promising representation for high-level spatial reasoning in autonomous mobile platforms. However, existing extraction frameworks typically rely on purely local visual clustering or str…

Scene Graph GenerationSpatial Reasoning

A Visual Semantic Adaptive Watermark grounded by Prefix-Tuning for Large Vision-Language Model

2026-01-12 · Qi Zheng, Shuliang Liu, Yu Huang, Sihang Jia 외 arxiv

Watermarking has emerged as a pivotal solution for content traceability and intellectual property protection in Large Vision-Language Models (LVLMs). However, vision-agnostic watermarks introduce visually irrelevant toke…

Visual Grounding