paper-with-me

홈 › Papers

Towards Tokenized Human Dynamics Representation

2021-11-22 · Kenneth Li, Xiao Sun, Zhirong Wu, Fangyun Wei, Stephen Lin

For human action understanding, a popular research direction is to analyze short video clips with unambiguous semantic content, such as jumping and drinking. However, methods for understanding short semantic actions cannot be directly translated to long human dynamics such as dancing, where it becomes challenging even to label the human movements semantically. Meanwhile, the natural language processing (NLP) community has made progress in solving a similar challenge of annotation scarcity by large-scale pre-training, which improves several downstream tasks with one model. In this work, we study how to segment and cluster videos into recurring temporal patterns in a self-supervised way, namely acton discovery, the main roadblock towards video tokenization. We propose a two-stage framework that first obtains a frame-wise representation by contrasting two augmented views of video frames conditioned on their temporal context. The frame-wise representations across a collection of videos are then clustered by K-means. Actons are then automatically extracted by forming a continuous motion sequence from frames within the same cluster. We evaluate the frame-wise representation learning step by Kendall's Tau and the lexicon building step by normalized mutual information and language entropy. We also study three applications of this tokenization: genre classification, action segmentation, and action composition. On the AIST++ and PKU-MMD datasets, actons bring significant performance improvements compared to several baselines.

📄 PDF Abstract BibTeX arXiv:2111.11433

Code (1)

likenneth/acton 공식 구현 pytorch

Tasks

Action SegmentationAction UnderstandingGenre classificationHuman DynamicsRepresentation Learning

Similar Papers 제목 키워드 기반

TVC: Tokenized Video Compression with Ultra-Low Bitrate

2025-04-22 · Lebin Zhou, Cihan Ruan, Nam Ling, Wei Wang 외

Tokenized visual representations have shown great promise in image compression, yet their extension to video remains underexplored due to the challenges posed by complex temporal dynamics and stringent bitrate constraint…

DecoderImage CompressionVideo Compression

Rethinking Tokenized Graph Transformers for Node Classification

2025-02-12 · Jinsong Chen, Chenyang Li, Gaichao Li, John E. Hopcroft 외

Node tokenized graph Transformers (GTs) have shown promising performance in node classification. The generation of token sequences is the key module in existing tokenized GTs which transforms the input graph into token s…

ClassificationNode ClassificationRepresentation Learning

DadaGP: A Dataset of Tokenized GuitarPro Songs for Sequence Models

2021-07-30 · Pedro Sarmento, Adarsh Kumar, CJ Carr, Zack Zukowski 외

Originating in the Renaissance and burgeoning in the digital era, tablatures are a commonly used music notation system which provides explicit representations of instrument fingerings rather than pitches. GuitarPro has e…

DecoderGenre classificationMusic GenerationMusic Style Transfer+1

TokenHMR: Advancing Human Mesh Recovery with a Tokenized Pose Representation

2024-04-25 · CVPR 2024 1 · Sai Kumar Dwivedi, Yu Sun, Priyanka Patel, Yao Feng 외

We address the problem of regressing 3D human pose and shape from a single image, with a focus on 3D accuracy. The current best methods leverage large datasets of 3D pseudo-ground-truth (p-GT) and 2D keypoints, leading t…

3D Human Pose EstimationHuman Mesh Recoveryvalid

Scalable Similarity Joins of Tokenized Strings

2019-03-21 · Ahmed Metwally, Chun-Heng Huang

This work tackles the problem of fuzzy joining of strings that naturally tokenize into meaningful substrings, e.g., full names. Tokenized-string joins have several established applications in the context of data integrat…

Data IntegrationFraud Detection