paper-with-me

홈 › Papers

A Study of Actor and Action Semantic Retention in Video Supervoxel Segmentation

2013-11-13 · Chenliang Xu, Richard F. Doell, Stephen José Hanson, Catherine Hanson, Jason J. Corso

Existing methods in the semantic computer vision community seem unable to deal with the explosion and richness of modern, open-source and social video content. Although sophisticated methods such as object detection or bag-of-words models have been well studied, they typically operate on low level features and ultimately suffer from either scalability issues or a lack of semantic meaning. On the other hand, video supervoxel segmentation has recently been established and applied to large scale data processing, which potentially serves as an intermediate representation to high level video semantic extraction. The supervoxels are rich decompositions of the video content: they capture object shape and motion well. However, it is not yet known if the supervoxel segmentation retains the semantics of the underlying video content. In this paper, we conduct a systematic study of how well the actor and action semantics are retained in video supervoxel segmentation. Our study has human observers watching supervoxel segmentation videos and trying to discriminate both actor (human or animal) and action (one of eight everyday actions). We gather and analyze a large set of 640 human perceptions over 96 videos in 3 different supervoxel scales. Furthermore, we conduct machine recognition experiments on a feature defined on supervoxel segmentation, called supervoxel shape context, which is inspired by the higher order processes in human perception. Our ultimate findings suggest that a significant amount of semantics have been well retained in the video supervoxel segmentation and can be used for further video analysis.

📄 PDF Abstract BibTeX arXiv:1311.3318

Code (0)

등록된 구현이 없습니다.

Tasks

object-detectionObject DetectionSegmentation

Similar Papers 제목 키워드 기반

Unveiling the Secrets of Engaging Conversations: Factors that Keep Users Hooked on Role-Playing Dialog Agents

2024-02-18 · Shuai Zhang, Yu Lu, Junwen Liu, JIA YU 외

With the growing humanlike nature of dialog agents, people are now engaging in extended conversations that can stretch from brief moments to substantial periods of time. Understanding the factors that contribute to susta…

DreamVideo: High-Fidelity Image-to-Video Generation with Image Retention and Text Guidance

2023-12-05 · Cong Wang, Jiaxi Gu, Panwen Hu, Songcen Xu 외

Image-to-video generation, which aims to generate a video starting from a given reference image, has drawn great attention. Existing methods try to extend pre-trained text-guided image diffusion models to image-guided vi…

Image to Video GenerationVideo Generation

JARViS: Detecting Actions in Video Using Unified Actor-Scene Context Relation Modeling

2024-08-07 · Seok Hwan Lee, Taein Son, Soo Won Seo, Jisong Kim 외

Video action detection (VAD) is a formidable vision task that involves the localization and classification of actions within the spatial and temporal dimensions of a video clip. Among the myriad VAD architectures, two-st…

Action DetectionRelationVideo Action Detection

BVB: Benchmarking Agentic Video Understanding via Programmatic Reconstruction in Blender

2026-09-14 · Yolo Y. Tang, Daiki Shimada, Jiayue Meng, Jing Bi 외 hf

Multimodal agents can create complex videos in software such as Blender by coding without relying on diffusion models. Yet video understanding benchmarks still evaluate models mainly through question answering. If an age…

Question Answering

Reinforcing User Retention in a Billion Scale Short Video Recommender System

2023-02-03 · Qingpeng Cai, Shuchang Liu, Xueliang Wang, Tianyou Zuo 외

Recently, short video platforms have achieved rapid user growth by recommending interesting content to users. The objective of the recommendation is to optimize user retention, thereby driving the growth of DAU (Daily Ac…

Recommendation Systemsreinforcement-learningReinforcement LearningReinforcement Learning (RL)