paper-with-me

홈 › Papers

Automatic Generation of Labeled Data for Video-Based Human Pose Analysis via NLP applied to YouTube Subtitles

2023-03-23 · Sebastian Dill, Susi Zhihan, Maurice Rohr, Maziar Sharbafi, Christoph Hoog Antink

With recent advancements in computer vision as well as machine learning (ML), video-based at-home exercise evaluation systems have become a popular topic of current research. However, performance depends heavily on the amount of available training data. Since labeled datasets specific to exercising are rare, we propose a method that makes use of the abundance of fitness videos available online. Specifically, we utilize the advantage that videos often not only show the exercises, but also provide language as an additional source of information. With push-ups as an example, we show that through the analysis of subtitle data using natural language processing (NLP), it is possible to create a labeled (irrelevant, relevant correct, relevant incorrect) dataset containing relevant information for pose analysis. In particular, we show that irrelevant clips ($n=332$) have significantly different joint visibility values compared to relevant clips ($n=298$). Inspecting cluster centroids also show different poses for the different classes.

📄 PDF Abstract BibTeX arXiv:2304.14489

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Learning-Based Human Segmentation and Velocity Estimation Using Automatic Labeled LiDAR Sequence for Training

2020-03-11 · Wonjik Kim, Masayuki Tanaka, Masatoshi Okutomi, Yoko SASAKI

In this paper, we propose an automatic labeled sequential data generation pipeline for human segmentation and velocity estimation with point clouds. Considering the impact of deep neural networks, state-of-the-art networ…

Segmentation

Discovery and recognition of motion primitives in human activities

2017-09-29 · Marta Sanzari, Valsamis Ntouskos, Fiora Pirri

We present a novel framework for the automatic discovery and recognition of motion primitives in videos of human activities. Given the 3D pose of a human in a video, human motion primitives are discovered by optimizing t…

Motion Generation

HACS: Human Action Clips and Segments Dataset for Recognition and Temporal Localization

2017-12-26 · ICCV 2019 10 · Hang Zhao, Antonio Torralba, Lorenzo Torresani, Zhicheng Yan

This paper presents a new large-scale dataset for recognition and temporal localization of human actions collected from Web videos. We refer to it as HACS (Human Action Clips and Segments). We leverage both consensus and…

Action ClassificationAction LocalizationAction RecognitionTemporal Action Localization+2

VideoScore: Building Automatic Metrics to Simulate Fine-grained Human Feedback for Video Generation

2024-06-21 · Xuan He, Dongfu Jiang, Ge Zhang, Max Ku 외

The recent years have witnessed great advances in video generation. However, the development of automatic video metrics is lagging significantly behind. None of the existing metric is able to provide reliable scores over…

Video GenerationVideo Quality Assessment

Lifting Unlabeled Internet-level Data for 3D Scene Understanding

2026-04-02 · Yixin Chen, Yaowei Zhang, Huangyue Yu, Junchao He 외 arxiv

Annotated 3D scene data is scarce and expensive to acquire, while abundant unlabeled videos are readily available on the internet. In this paper, we demonstrate that carefully designed data engines can leverage web-curat…

Visual Question AnsweringInstance Segmentation3D Object DetectionScene Understanding